Tech News
qm – Multiplayer agent harness for work
qm is a multiplayer agent harness for work, featuring an 'anti-slop' frontend skill that enforces design quality and avoids templated outputs, with support for MCP clients and multi-agent collaboration.
DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
DeepSeek V4 Flash 0731 is a new frontier model with strong price-performance, evaluated on coding agent tasks using a minimal mode of DeepSeek Harness, and is noted as a daily driver for cost-effective coding.
Co-Designing AI Model Attention for Fast, Interactive Long-Context Inference
As agentic and long-context workloads become common, the context lengths increase and attention consumes a larger share of inference time.
Announcing the Agentic Catalog Experience in Amazon Quick
Amazon Quick introduces the Agentic Catalog Experience, an AI-powered workflow for data curators to discover upstream catalog assets in natural language and auto-create Datasets and Topics with inherited semantics.
How enabling two settings tripled our scores on the ARC-AGI-3 benchmark
How two API settings improved GPT-5.6 performance on ARC-AGI-3, boosting scores and efficiency by retaining reasoning and enabling compaction.
GitHub Repos
sickn33/agentic-awesome-skills
AAS Core is the local, agent-first control plane for complete catalog discovery, agent-owned selection, stack validation, and planning, backed by 1,987+ agentic skills.
# Python# agent skills# agentic skillsOpenHands/OpenHands
🙌 OpenHands: AI-Driven Development
# TypeScript# agent# artificial intelligencenitrocloudofficial/nitrostack
The full-stack TypeScript framework to build, test, and deploy production-ready MCP servers and AI-native apps.
# TypeScript# agentic ai# aiResearch Papers
VidMap: Exploiting Temporal Structure for Video-Based Structure-from-Motion
VidMap is a new system that merges the strengths of SLAM and Structure-from-Motion to reconstruct 3D scenes from ordinary videos, even without camera calibration. It uses the video's natural order to improve accuracy and robustness, making it easier to create 3D models from...
Would You Walk to the Car Wash? Revealing the Salience Bias of Large Language Models in Commonsense Reasoning
This paper introduces SaliTrap, a benchmark showing that LLMs often ignore common-sense constraints when given distracting explicit details (like numbers), and that this is mostly because the knowledge is suppressed by the distracting context, not because the model lacks the...
Beyond a Single Judge: Simulating Social Persona Panels for Generative UI Evaluation
The Evidence-Grounded, Social-Weighted Persona Panel is proposed, a three-stage GenUI evaluation method in which a panel of psychologically diverse, evidence-grounded personas independently rates a screenshot, exchanges opinions under a trait-derived, semantically-gated...
