Tech News
Felony Bench
is a benchmark that counts unique instances where AI agents inadvertently compromise or affect third-party entities, raising questions about legal accountability for agentic AI actions under laws like CFAA.
llm-openrouter 0.7
Now that this plugin is compatible with LLM 0.32 it works much better with reasoning LLMs available through OpenRouter. Updated for compatibility with LLM.
Agentic Data Operations Platform (ADOP): Data engineering into hours
The Agentic Data Operations Platform (ADOP) is a reference architecture on Amazon Bedrock that uses specialized AI agents to automate the full Bronze-to-Silver-to-Gold data pipeline lifecycle, compressing new-source...
NVIDIA AVO Reaches 100% on ARC-AGI-3, Demonstrating a Frontier-Level General-Purpose Architecture for Long-Horizon Autonomous Agents
A frontier language model is only one component of an AI agent. The surrounding agent system—often called a harness—determines how the model receives.
DeepSeek-v4-flash-vision-exp
DeepSeek's new vision model, v4-flash-vision-exp, shows promise in handling screenshots but still struggles with precise visual tasks like reading clocks, as noted by developers.
GitHub Repos
ComposioHQ/composio
Composio powers 1000+ toolkits, tool search, context management, authentication, and a sandboxed workbench to help you build AI agents that turn intent into action.
# TypeScript# agentic ai# agentsComposioHQ/awesome-claude-skills
A curated list of awesome Claude Skills, resources, and tools for customizing Claude AI workflows
# Python# agent skills# ai agentsvercel/workflow
Workflow SDK: Build durable, reliable, and observable apps and AI Agents in TypeScript
# TypeScript# ai agents# durable executionResearch Papers
Large Models for Small Devices: Recent Advances and Empirical Analysis of Edge AI Deployment
This paper surveys and tests various model compression techniques (like pruning and quantization) for running AI models on small devices, finding that no single method works best and that compression can sometimes hurt performance or even make models look better than they are.
ATLAS: Scaffold-Free Algorithm Synthesis by LLMs via Embedding-Guided Quality-Diversity Search
ATLAS is a new method that uses large language models and quality-diversity search to automatically design complete algorithms for combinatorial optimization problems without needing a predefined scaffold. It shows promising results on four NP-hard problems, producing diverse...
A Declarative-Procedural Perspective on Expert Routing in Bilingual Mixture-of-Experts Language Models
This paper explores whether bilingual AI language models organize their internal 'expert' pathways in a linguistically meaningful way, similar to how humans separate grammar and vocabulary. The authors found that training order affects this organization, with mixed-language...
