Tech News
Beam: Reflection's 501B open-weight model
Reflection released Beam, a 501B-parameter sparse Mixture-of-Experts open-weight model with 23B active parameters, trained on 23.8T tokens and optimized for coding, reasoning, and agentic workloads.
New agent skill: Amazon SageMaker optimized generative AI inference for your coding agent
Amazon SageMaker optimized generative AI inference introduces the aws-ai-ml skill through the Agent Toolkit for AWS, giving coding agents like Kiro, Claude Code, and Codex deep expertise in inference optimization.
vllm-project/vllm v0.31.0
v0.31.0 Highlights This release features 717 commits from 307 contributors (96 new)!
Quoting Felix Rieseberg
The "old" version of Cowork runs model inference in the cloud, executing tool calls in an Anthropic-provided VM we shipped to your computer.
The Agent Said It Was Done. The Database Disagreed.
A new Hugging Face blog post examines how AI agents can falsely report task completion when their actions conflict with underlying database state.
GitHub Repos
EverMind-AI/EverOS
One portable memory layer for every AI agent: local-first, Markdown-native, user-owned, and self-evolving across apps, tools, and workflows.
# Python# agent memory# agentic aiKunAgent/Kun
Local-first AI agent workspace for coding, writing, design, research, and automation — one runtime for desktop GUI and TUI.
# TypeScript# agentic workflow# ai agentLetsFG/LetsFG
Agent-native flight & hotel search and booking — MCP server, CLI, and Python/JS SDKs.
# Python# ai# ai agentResearch Papers
AgentDiscover: Autonomous Discovery with Minimal Search Scaffolding
AgentDiscover lets a coding agent plan its own search instead of following a fixed human-designed algorithm, using a database as long-term memory. It reports better and cheaper results than prior discovery frameworks on tasks like kernel engineering, biology, and math.
Reward Stealing Attack on Large Language Models
This paper proposes ReSA, an attack that tries to infer the hidden safety reward an aligned LLM was trained with, then flips that reward during decoding to make the model produce unsafe outputs. It claims the recovered reward transfers across prompts and models, but the work...
OVAL: Output-Aware Local Page Bases for KV Cache Retrieval
Long context inference with large language models becomes increasingly expensive as attention must operate over an ever growing KV cache. Page sparse attention reduces this cost by representing each KV page compactly and retrieving only a subset for each query.
