Tech News
Revealing the details of how OpenAI agents hacked Hugging Face
Public traces reveal how OpenAI agents pursued a goal in an open-ended environment, exposing failure modes like brute-force search, weak consolidation, and noisy high-volume requests.
Quoting John Gruber
Muse is getting a lot of attention — including mine — because it’s both groundbreaking technically (each user gets their own entire persistent Linux VM running in Meta’s cloud) and because it’s packaged.
From portal-hopping to instant answers: HEMA’s journey with MCP and Amazon Bedrock
HEMA, a 100-year-old Dutch retailer, turned developer portal-hopping into instant answers by building HAL, an internal AI assistant on Amazon Bedrock AgentCore.
Ringg’s AI agents resolve up to 65% of customer calls with OpenAI
Using GPT-5.6, Ringg powers multilingual agents across voice, chat, WhatsApp, and web for 90% less cost.
How SWE-Serve Exposes the Gap Between Local Tests and Live Serving
An AI coding agent’s patch can pass tests yet fail when the server loads a real model and handles requests. Evaluating changes to inference-serving software.
GitHub Repos
unclecode/crawl4ai
Open-source web crawler and scraper for LLMs and AI agents: any website into clean, LLM-ready Markdown.
# Python# ai# ai agentsruvnet/ruflo
🌊 The original agent harness.
# TypeScript# agentic ai# agentic frameworkbeenuar/AiSOC
Open-source AI Security Operations Center: alert fusion, LLM-agent triage, MITRE ATT&CK investigation, and a replayable decision ledger for every agent step.
# Python# agentic ai# ai securityResearch Papers
Back to the Definition: Estimating Step-Level Advantages via Trajectory Graphs for Agentic Reinforcement Learning
This paper proposes GRAFT, a method to estimate step-level advantages in multi-turn agentic reinforcement learning by building a graph of rollout trajectories and using Bellman iteration to assign credit to each step. It aims to fix the bias in group-based methods like GRPO...
SciWalker: Synthesizing Scientific Coding Problems with Operator Graphs and Execution Feedback
SciWalker automatically creates scientific coding problems by combining library operations into graphs, sampling workflows, and using execution feedback to fix generated problems. It builds 8,178 problems and shows that training on them improves a model's scientific coding...
SAGE: Mitigating Long-Horizon Reasoning Biases via Topological Guidance
This work proposes SAGE (Structural Admissibility-Guided Exploration), a unified framework that injects structural guidance to alleviate exploration bias and compounding bias in long-horizon reasoning and achieves up to an 8-fold improvement on the Andrews-Curtis problem.
