Tech News
Handbook.md shows that long policy documents do not reliably govern agents
A new benchmark reveals that LLMs fail to reliably follow long policy documents, undermining claims of superhuman reasoning and highlighting fundamental limits in context handling.
AI Worming through Word
Neat new prompt injection variant by Håkon Måløy, who found a way to upgrade prompt injection attacks against Microsoft Word to full self-replicating worms: An attacker places hidden instructions in a document.
How enabling two settings tripled our scores on the ARC-AGI-3 benchmark
How two API settings improved GPT-5.6 performance on ARC-AGI-3, boosting scores and efficiency by retaining reasoning and enabling compaction.
Generate Autonomous Business Insights with AI Agent and MCP Servers
Learn how Amazon Bedrock AgentCore delivers autonomous, cross-system business intelligence through configuration rather than custom code.
Six Agent Harness Capabilities for Higher Model Performance
Building a great AI agent isn’t just about choosing the right models. The harness is the architecture surrounding the model. How it renders context, executes.
GitHub Repos
hesreallyhim/awesome-claude-code
A hand-picked collection of the finest of resources for the most awesome of agents, Claude Code, the undisputed champion of coding companions, from the unstoppable team at...
# Python# agent skills# agentic codecode-yeongyu/oh-my-openagent
omo/lazycodex: The coding agent for tokenmaxxers;the one and only agent harness for complex codebases.
# TypeScript# ai# ai agentsjgravelle/jcodemunch-mcp
Cut AI token costs 95%+ on code exploration.
# Python# ai coding# ai toolsResearch Papers
Hierarchical Spatio-Temporal Transformer for Coherent Emergency Department Forecasting
This paper proposes a hierarchical Transformer model for forecasting emergency department demand at hospital, regional, and national levels simultaneously, ensuring coherence across levels, and demonstrates significant improvements on a new Portuguese dataset.
SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response
This paper presents SecRespond, the first benchmark to evaluate LLM agents on post-compromise incident response tasks using forensic disk snapshots and security alerts, finding that current agents can handle alerts but fail at proactive investigation and remediation.
VidMap: Exploiting Temporal Structure for Video-Based Structure-from-Motion
This paper introduces VidMap, a system that combines the strengths of SLAM and Structure-from-Motion to recover camera poses and 3D structure from arbitrary, long, uncalibrated videos. It uses temporal ordering and global optimization to handle challenging motions and visual...
