Tech News
METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
A postmortem of the HuggingFace hack by METR and Redwood, discussed on Hacker News, examines the incident's implications for AI agent safety and institutional oversight.
Understanding ChatGPT Work
OpenAI announced ChatGPT Work on July 9th, and have been furiously iterating on it ever since. It is an extraordinarily confusing and very powerful product. Here's what I've figured out about it so far.
Build agentic creative workflows with Amazon Quick and fal
Creative teams produce more assets than ever, but fragmented tools and manual context transfer slow production.
vllm-project/vllm v0.28.0
v0.28.0 Highlights This release features 584 commits from 270 contributors (76 new)!
“I just chose words carefully”
A developer's reflection on deliberately choosing words in code and documentation sparks a discussion on how precise language shapes readability and design, drawing parallels to scriptwriting and UI naming.
GitHub Repos
sickn33/agentic-awesome-skills
AAS Core is the local, agent-first control plane for complete catalog discovery, agent-owned selection, stack validation, and planning, backed by 2,005+ agentic skills.
# Python# agent skills# agentic skillsopendatalab/MinerU
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
# Python# ai4science# document analysisusewhale/Whale
Whale — blazingly fast, terminal-first AI coding agent for DeepSeek.
# Go# coding agent# deepseekResearch Papers
ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL
ContextPilot improves how AI agents manage their working memory during long tasks by adding planning, long-term memory, and compression tools, and uses a new reinforcement learning method to better assign credit to context-editing actions.
Blind Men and the Elephant: Probing the Epistemic Myopia of LLMs under Long-Tail Divergent Knowledge
This paper presents ElephantBench, a benchmark of 1,094 questions about long-tail facts with multiple valid answers, and shows that even the best LLMs often recall only one account, missing others.
JANUS: Online Jacobian-Aligned Infill for Black-Box Optimization
JANUS is a plug-and-play infill module that extracts a local Jacobian from the recent evaluation trace, and gives the best mean cost on 1135-dimensional UAV path planning and improves SMS-EMOA/AGE-MOEA2 hosts on 12/38 multi-objective tasks with zero significant regressions.
