Tech News
Kolibri: A Sovereign Open-Weight Model
Aleph Alpha released Kolibri, an open-weight agentic LLM with a detailed technical report that serves as a tutorial for building modern agentic models, including dataset construction and training with abstention data...
A model guide for the GPT-6 family
Learn how startups can choose GPT-6 models, tune reasoning effort, improve prompts and skills, coordinate tools, and prepare workflows for production.
We're going to need default hard budget caps on pretty much everything
Here's a product feature which the world is going to need a whole lot more of over the coming months and years: default hard budget caps.
Sweep thousands of leases for compliance using Amazon Quick and the Adjudicated Query pattern
The Adjudicated Query pattern pairs the Amazon Quick chat agent with a bounded MCP server over a deterministic rules engine to deliver provably complete, defensible compliance answers.
modelcontextprotocol/python-sdk v2.3.0
pip install -U mcp. Docs: Mostly fixes, plus three new options. A few things behave differently, so skim these first: Behaviour changes httpx2>=2.10.0 is now required (#3600) It was.
GitHub Repos
calesthio/OpenMontage
World's first open-source, agentic video production system.
# Python# agent# agentic aicallstack/agent-device
Mobile app automation and verification for AI coding agents.
# TypeScript# adb# agentic aiUnicomAI/wanwu
China Unicom's Yuanjing Wanwu Agent Platform is an enterprise-grade, multi-tenant AI agent development platform.
# Go# agent# agentic aiResearch Papers
HazardWeaver: Scientific Route Selection for Hazard Analysis Agents
This paper builds an AI agent that chooses which scientific methods to use for hazard analysis, updating its choices as new data arrives. It also introduces a benchmark to test both the answers and the decision paths.
UniIntervene++: An Adaptive Intervention Agent for Efficient Real-World Reinforcement Learning
UniIntervene++ is a method for robot reinforcement learning that decides when and how a human should help, adapting the amount of help as the robot gets better. It combines the robot's own policy, human corrections, and a task-specific code policy into one decision framework...
A Near-Zero Monitor Readout Is Not Evidence of Behavioral Control
Post-training with verifiable rewards can induce reward hacking, motivating the use of monitors within the training objective rather than solely for offline auditing. We show that a low monitor readout does not identify whether such an intervention controls behavior.
