← Archive / All Digests
A wolf in round glasses reading a book, wrapped in a golden ribbon, in a sunlit forest.

Wolf Digest — Tuesday, July 21, 2026

Coverage window: 2026-07-20 03:41 ET2026-07-21 03:02 ET
Press play to listen
Tuesday, July 21, 2026
10m 26s · top-4 narrated briefing
#1 · Industry
China's open-weight surge forces the US-China AI question into the open
The debate over how far Chinese open-weight models have closed the gap with the American frontier reached a head this week, and it was the dominant AI story across nearly every source. The proximate cause is Moonshot's Kimi K3, a 2.8-trillion-parameter mixture-of-experts model re…
8.9 · 7 srcs
#2 · Infrastructure
Google is said to be building a 'Frozen' chip with Gemini baked into the silicon
Google is developing a server chip that integrates the blueprint of its Gemini model directly into the hardware, according to The Information, which reports the part is informally called "Frozen v2." Instead of running a model as software on a general-purpose accelerator, the des…
7.7 · 2 srcs
#3 · Government & Defense
Anduril unveils Thunder, an autonomous tiltrotor 'loyal wingman' for helicopters
Anduril used the Farnborough Air Show to unveil Thunder, an autonomous attack rotorcraft prototype built to fly alongside crewed military helicopters. Described as a Group 5 unmanned tiltrotor, Thunder is designed to function as a loyal wingman for rotary-wing aircraft in the sam…
7.7 · 2 srcs
6.5
#1
Industry 2026-07-20 Hacker NewsThe InformationInterconnects (Nathan Lambert)MIT Technology ReviewStratecheryTechCrunchImport AI (Jack Clark) 8.9 8.4/8.6/9.6

The debate over how far Chinese open-weight models have closed the gap with the American frontier reached a head this week, and it was the dominant AI story across nearly every source. The proximate cause is Moonshot's Kimi K3, a 2.8-trillion-parameter mixture-of-experts model released the prior Thursday with weights promised for July 27. It is the largest open-source model in the world, and it did not just post respectable numbers: it topped a key coding leaderboard from Arena, placing ahead of both OpenAI's GPT-5.6 and Anthropic's most capable Claude Fable 5. Arena chief executive Anastasios Angelopoulos argued the result undercuts the comfortable assumption that Chinese labs advance only by distilling American models, saying it is the first time the narrative that they "might actually just be really good at developing models" has broken through.

Nathan Lambert's read is that the open-to-closed and American-to-Chinese gaps have compressed from a debated six-to-nine months down to roughly three-to-five, and that K3 is the closest open weights have come to the frontier since DeepSeek R1. Where R1 was a fast pivot to reasoning, K3 is a Chinese lab executing cleanly on known scaling. The UK's AI Security Institute, summarized in Jack Clark's Import AI, put concrete numbers on a narrower slice: on seventy cyber evaluations, GLM-5.2 now performs like Claude Opus 4.6 did four-and-a-third months earlier, and DeepSeek-V4-Pro lands between last year's Claude Opus 4.5 and GPT-5, a tighter lag than the six-to-ten months measured through most of 2025.

The commercial logic is straightforward and is what makes this more than a benchmark story: every capable free model out of China gives enterprises another reason not to pay for metered access to OpenAI or Anthropic. That pressure has spilled into policy. Reporting from The Information notes the administration had discussed banning open-source models last year, and that K3's arrival revived those conversations; advisors and current officials aired sharp public disagreements over the right response. Moonshot, meanwhile, is seeking investor approval to begin a Hong Kong initial public offering. A countervailing signal comes from Beijing: officials have held closed-door talks since June about restricting foreign access to China's most advanced models, open and closed alike, which would complicate any assumption that the next Alibaba, Zhipu, or DeepSeek release ships globally with downloadable weights on day one.

The through-line is that open weights have moved from a cost-saving curiosity to a strategic variable that touches model economics, export policy on both sides of the Pacific, and the business models of the largest American labs at once.

How it was discussed
  • Arena's Angelopoulos frames K3 as evidence Chinese labs build, not just distill, topping a coding leaderboard over GPT-5.6 and Fable 5.
  • Interconnects quantifies the compression: frontier gap down to three-to-five months, the closest open weights since DeepSeek R1.
  • Import AI cites UK AISI cyber-eval numbers showing a four-to-seven-month lag, narrower than 2025's six-to-ten.
  • The Information and MIT Technology Review focus on Washington's split response and revived talk of restricting foreign open models.
  • Gradient Flow flags the mirror-image risk: Beijing itself weighing export limits on its most capable models.
open-weights Kimi K3 US-China
#2
Infrastructure 2026-07-20 The InformationTechCrunch 7.7 8.0/7.9/7.2

Google is developing a server chip that integrates the blueprint of its Gemini model directly into the hardware, according to The Information, which reports the part is informally called "Frozen v2." Instead of running a model as software on a general-purpose accelerator, the design etches the model's structure into the silicon itself, trading flexibility for efficiency. Engineers working on the chip project it could be six to ten times more efficient than the newest generation of Google's homegrown TPUs, measured as tokens served per unit of power, when it launches.

The motivation is a capacity crunch. The Information reports the shortage of AI computing has fueled internal tension and pushed Google Cloud to turn down deals with outside customers, so a large efficiency gain on inference translates directly into more served demand from the same power and floorspace envelope. Investors read it as material: Alphabet stock rose about three percent on Monday after the report, and TechCrunch corroborated that Google is working on a chip designed to make Gemini run much more efficiently.

The engineering bet is the interesting part. Baking a specific model into hardware only pays off if that model's architecture is stable enough to justify a multi-year silicon cycle, which is why the "frozen" framing matters: it presumes the shape of a frontier model changes slowly enough that specialization beats generality. If that holds, tokens-per-watt on a model-specific part could become a durable cost advantage over renting general accelerators, and it sharpens the divergence between labs that own their stack end to end and those that buy compute on the open market. If model architectures keep shifting underneath, the same rigidity becomes a liability. Either way, the reported six-to-ten-times figure is large enough that, if it survives contact with production, it reframes the efficiency conversation around whether the model or the accelerator is the right thing to hold fixed.

How it was discussed
  • The Information reports the six-to-ten-times tokens-per-watt projection and ties it to a capacity crunch forcing Google Cloud to refuse customers.
  • TechCrunch corroborates the effort and its Gemini-efficiency aim; markets moved Alphabet up about three percent on the news.
Google custom silicon TPU Gemini
#3
Government & Defense 2026-07-20 DefenseScoopDefense One 7.7 7.1/6.6/6.5 +1.0 gov_defense

Anduril used the Farnborough Air Show to unveil Thunder, an autonomous attack rotorcraft prototype built to fly alongside crewed military helicopters. Described as a Group 5 unmanned tiltrotor, Thunder is designed to function as a loyal wingman for rotary-wing aircraft in the same way Collaborative Combat Aircraft are being developed to accompany fighter jets. The company says it has been working on the platform for roughly three years and developed it together with Archer Aviation, drawing on Archer's tiltrotor propulsion experience.

Anduril President Chris Brose framed the design around maneuver warfare and the tactical, air-ground littoral fights seen in Ukraine. The stated goal is to bring mass to a contested deep fight, where crewed helicopters are increasingly exposed and where an attritable autonomous escort can absorb risk, extend sensing, and add magazine depth without putting more crews forward. Thunder is aimed at future Army, Marine Corps, and Special Operations Command missions, and it extends the loyal-wingman concept from fixed-wing formations to the rotary-wing environment, where lower altitudes and shorter engagement ranges impose different autonomy and survivability requirements.

The unveiling lands amid a broader push toward manned-unmanned teaming across the services, and it positions Anduril, better known for its software-defined systems and smaller drones, as a prime pursuing a large, complex air vehicle. The open questions are the ones that always attend a prototype: how much of the autonomy stack is flight-proven versus aspirational, what the unit cost looks like at the "mass" the company invokes, and how the tiltrotor configuration trades speed and range against the hover and low-speed agility that helicopter escort demands. Defense One's coverage frames the same system explicitly as a robot wingman pitched to the Army, underscoring that the near-term customer conversation is about augmenting existing helicopter formations rather than replacing them.

How it was discussed
  • DefenseScoop details the Group 5 tiltrotor design, the Archer Aviation partnership, and Brose's maneuver-warfare and Ukraine framing.
  • Defense One frames Thunder as a robot wingman pitched specifically to US Army helicopter formations.
Anduril loyal wingman Farnborough autonomy
#4
Industry 2026-07-20 TechCrunch 7.5 7.4/8.0/7.1

A federal judge gave final approval on Monday to Anthropic's $1.5 billion settlement of a class-action copyright suit brought by authors and publishers, clearing the company to begin paying claimants. Judge William Alsup of the Northern District of California had granted preliminary approval last year; after his retirement, Judge Araceli Martinez-Olguin signed off on the final deal. The payout is structured as roughly three thousand dollars per work across an estimated five hundred thousand works, split among the rights holders, and is believed to be the largest recovery in the history of US copyright law.

Many authors still do not regard it as a win, and the reason is how the underlying legal question was resolved. Alsup sided with Anthropic on the central issue, ruling that training a model on copyrighted text is fair use, a decision widely read as a turning point for the industry. What the ruling did not excuse was how Anthropic obtained the books. The company had assembled its training library from two streams: volumes it purchased and scanned, which the court treated as permissible, and books downloaded from pirate repositories such as Library Genesis and Pirate Library Mirror, which Alsup found unlawful on their own terms. He signaled the piracy question could go to a jury, and Anthropic settled soon after to avoid a trial and the open-ended damages a verdict might have carried.

Because the settlement resolves a single district-court case and forecloses an appeal, it does not create binding precedent. Other judges remain free to reach different conclusions on their own records, and several are positioned to do so. Active copyright suits continue against Google, Meta, Midjourney, and OpenAI over training on protected works, and only last week a group including Hachette, Cengage, Elsevier, and the author Scott Turow filed a fresh class action against Google over the use of their material to train Gemini. The practical takeaway for model builders is narrow but sharp: the fair-use theory for training survived, but data provenance is now the exposed flank, and how a corpus was acquired can be as consequential as what was in it.

Anthropic copyright fair use litigation
#5
Safety, Policy & Regulation 2026-07-20 Import AI (Jack Clark) 6.9 7.0/7.4/6.4

The UK AI Security Institute published its first analysis of how far leading open-weight models trail the closed cyber frontier, and the gap has narrowed. Across seventy narrow cyber-capability evaluations, GLM-5.2 performs closest to Claude Opus 4.6, released about 4.3 months earlier, while DeepSeek-V4-Pro sits between Claude Opus 4.5 and GPT-5. AISI frames this as a compression from the six-to-ten-month lag measured through most of 2025 to roughly four-to-seven now, giving a concrete, security-relevant measurement to the broader open-versus-closed convergence story.

cyber evals open-weights AISI
#6
Industry 2026-07-20 Hacker NewsNikkei Asia 6.9 6.6/7.6/6.6

A Nikkei analysis circulating widely on Hacker News, at 186 points, tallies roughly $1.65 trillion in hidden or off-balance-sheet liabilities across five large US technology firms, driven by the increasingly opaque financing structures behind AI data-center and compute buildouts. The piece argues that special-purpose vehicles, leasing arrangements, and vendor financing obscure the true leverage funding the current capex wave, making the sector's balance-sheet risk harder to read than headline cash positions suggest. It is a systemic-finance counterpoint to the capability-race coverage dominating the day.

infrastructure capex finance
#7
Safety, Policy & Regulation 2026-07-20 OpenAI ResearchHacker News 6.8 7.0/7.5/6.0

OpenAI published a research note on safety and alignment for long-horizon models, systems that run for extended periods and take many chained actions rather than answering a single prompt. The post shares lessons from deploying such models, describing new categories of risk that emerge over long rollouts, specific observed failures, and safeguards refined through iterative deployment. The framing reflects the shift from single-turn chat safety to agentic settings, where errors compound across steps and where monitoring, interruption, and recovery become first-class safety mechanisms rather than afterthoughts.

How it was discussed
  • The OpenAI post frames long-running agents as a distinct safety regime; the Hacker News thread centered on how much of the mitigation is procedural versus technical.
safety agents long-horizon
#8
AI Coding 2026-07-20 Hacker NewsCursor (Anysphere) 6.7 7.0/6.5/6.6

Cursor's engineering blog, at 176 points on Hacker News, argues that running many coding agents in parallel, "agent swarms," changes the cost structure of AI-assisted development. When dozens of agents attack a task concurrently, aggregate token consumption dwarfs single-agent workflows, shifting the binding constraint from model quality to throughput, cost per token, and orchestration overhead. The post makes the case that cheaper, faster models plus good harness design can outperform a single call to a stronger, pricier model, which reframes model selection as a portfolio-and-scheduling problem rather than a quality ranking.

agents coding economics Cursor
#9
Government & Defense 2026-07-20 Shield AI 6.7 5.9/5.7/5.4 +1.0 gov_defense

Shield AI and GE Aerospace completed integration and engine light-off of an Axisymmetric Vectoring Exhaust Nozzle on GE's F110-GE-129E engine for Shield AI's X-BAT aircraft, a step the companies call a critical milestone toward vertical flight. X-BAT is designed to take off and land vertically without a runway, and the vectoring nozzle is what makes that possible on a platform of its size and thrust class. The companies describe it as the first fully integrated test campaign of its kind since the 1990s, bringing nozzle hardware, controls, and engine systems together at GE's Peebles, Ohio test site.

Shield AI VTOL propulsion
#10
Government & Defense 2026-07-20 C4ISRNET 6.5 5.6/5.5/5.4 +1.0 gov_defense

Lockheed Martin unveiled Morfius X-Rotor, a counter-drone system the company says is built to disable large numbers of hostile drones in a single sortie, part of a wave of counter-swarm hardware responding to the proliferation of cheap uncrewed threats. The reveal sits alongside other counter-UAS announcements this cycle, including Israel Aerospace Industries' "Hypnosis" system for jamming and spoofing the satellite navigation that guides attacking drone swarms, reflecting how central electronic and kinetic counter-drone capability has become to air-defense procurement.

counter-drone Lockheed air defense
#11
Research 2026-07-20 Hacker News 6.5 6.4/6.6/6.4

A widely discussed analysis, 210 points on Hacker News, attempts to measure the prevalence of AI-assisted writing across arXiv submissions over time and is candid about where the measurement breaks down. Detectors keyed on stylistic markers and characteristic phrasing conflate genuine LLM authorship with light editing, non-native English, and template reuse, so absolute prevalence estimates are fragile even where relative trends look real. The piece is a useful caution about over-reading AI-detection statistics in scientific-writing debates.

AI writing measurement arXiv
#12
AI Coding 2026-07-20 Hugging Face Daily PapersAK (@_akhaliq) Daily PapersarXiv cs.CLarXiv — Evals & Benchmarks 6.4 6.6/6.3/6.4

SWE-Pruner Pro tackles long-context management for coding agents by dropping the separate code classifier that prior methods such as SWE-Pruner bolt on. The authors show the agent model itself already carries the signal for which files and spans are irrelevant to the current task, so pruning can be driven from the model's own internal judgments rather than an external module. The result is cheaper, better-aligned context compression for software-engineering agents, where context bloat is a dominant cost driver.

cs.CL agents context pruning
#13
Reinforcement Learning 2026-07-20 Hugging Face Daily PapersAK (@_akhaliq) Daily PapersarXiv cs.CLarXiv — EfficiencyarXiv — Reinforcement Learning 6.4 6.6/6.4/6.2

Reinforcement learning on open-ended tasks compresses rich rubric-based evaluation into a single scalar reward, discarding textual feedback and conflating responses with very different quality profiles. LLM-as-a-Coach proposes an experiential-learning scheme that retains the natural-language critique as a learning signal, letting the model condition on why a response fell short rather than only how much. Across non-verifiable tasks, the approach improves over scalar-reward baselines by preserving information the reward function throws away.

RL reward modeling rubrics
#14
Robotic Autonomy 2026-07-20 Hugging Face Daily PapersAK (@_akhaliq) Daily PapersarXiv cs.RO 6.4 6.5/6.4/6.2

RynnBrain 1.1 is a family of embodied foundation models spanning 2B, 9B, and a 122B-A10B mixture-of-experts scale, trained in a unified spatio-temporal and physically grounded framework. The models support embodied perception, spatial reasoning, and action across scales, with the MoE variant meant to hold capability while limiting active parameters at inference. It is another entry in the fast-growing vision-language-action foundation-model line, and the multi-scale release is aimed at spanning from on-robot deployment to server-side reasoning.

VLA embodied MoE
#15
Industry 2026-07-20 NVIDIA AI Blog 6.3 6.2/6.2/6.4

At SIGGRAPH, Nvidia laid out graphics and simulation updates centered on agentic and physical AI, tying its rendering and world-simulation stack to robotics and embodied workflows. The company also highlighted a life-science deployment with Bristol Myers Squibb built around a large on-premise AI installation. The throughline is Nvidia's continued positioning of simulation as the substrate for training embodied agents, connecting graphics tooling to the robot-learning and generative-media pipelines that increasingly depend on synthetic data.

Nvidia SIGGRAPH physical AI
#16
Agents & Tool Use 2026-07-20 TechCrunch 6.3 6.2/6.4/6.2

The Model Context Protocol, the increasingly standard interface for connecting models to external tools and data, is adopting a looser "stateless" approach to session identifiers on the server side, closer to how ordinary web services handle sessions. The change lowers the implementation burden for server authors, who no longer need to maintain sticky per-session state, and should make MCP servers easier to build and scale. It is an incremental but meaningful ergonomics improvement for the plumbing underneath most current tool-using agent stacks.

MCP agents protocol tools
#17
AI Coding 2026-07-20 Hacker News 6.3 6.3/6.0/6.6

Nativ, a tool for running frontier open-weight models locally on Apple silicon, drew 250 points and active discussion on Hacker News. Interest tracks the day's larger open-weights theme: as capable open models proliferate, local-inference tooling that squeezes them onto consumer Macs lowers the barrier to running strong models without cloud dependence or per-token cost. The thread centered on achievable throughput and memory footprint on high-unified-memory Apple machines.

local inference Apple silicon open-weights
#18
Safety, Policy & Regulation 2026-07-20 Gradient Flow (Ben Lorica) 6.3 6.3/6.6/6.0

Since June, Chinese officials have reportedly held closed-door talks about restricting foreign access to the country's most advanced AI models, both open and closed, with the Ministry of Commerce leading and firms including Alibaba, ByteDance, and Zhipu involved. Nothing is decided, and options reportedly range from light registration to keeping the most capable future models domestic-only. The practical implication for builders is concrete: assuming every future Chinese release ships globally with downloadable weights on day one is no longer safe, and single-model-family dependencies now carry supply risk.

China export controls open-weights
#19
Reinforcement Learning 2026-07-20 arXiv cs.AIarXiv cs.CLarXiv — Evals & BenchmarksarXiv — Reinforcement Learning 6.3 6.4/6.3/6.1

MADA-RL targets the high training cost of reasoning RL for compact models at or below four billion parameters. It folds multi-agent debate structure into a parameter-efficient RL objective, using debate-style disagreement among rollouts to shape the reward signal, so small models get more informative gradients under tight compute budgets. The paper reports reasoning gains over standard RL baselines at compact scale, where the usual recipe is prohibitively expensive.

RL multi-agent debate small models
#20
Evaluations & Benchmarks 2026-07-20 arXiv cs.AIarXiv — Evals & BenchmarksarXiv — Post-trainingarXiv — Reinforcement Learning 6.3 6.4/6.4/6.1

LLM judges increasingly score whether clinical models are overconfident under incomplete evidence, but this work asks whether a measured "safety gain" reflects genuine behavior change or merely the judge's calibration. The authors find safety improvements can be judge-dependent while helpfulness costs are model-specific, meaning a headline safety number can move without the underlying model becoming safer. It is a pointed methodological caution for the growing practice of using LLM judges to certify medical-model safety.

evals LLM-as-judge clinical
#21
AI for Science 2026-07-20 arXiv — AI for SciencearXiv cs.AIarXiv — Evals & BenchmarksarXiv — Generative Media 6.3 6.4/6.3/6.1

Generative models for materials design explore immense chemical spaces, but a large share of AI-proposed compositions are physically implausible, violating established chemical rules. This paper introduces fast chemical filters for ultra-high-throughput screening and generation that discard invalid candidates cheaply before expensive validation, raising the yield of plausible structures per unit compute. It is a practical throughput lever for generative materials pipelines, where the bottleneck is separating the credible minority from a flood of invalid samples.

materials generative screening
#22
Safety, Policy & Regulation 2026-07-20 TechCrunchThe Information 6.2 6.0/6.4/6.2

Chris Fall, director of the Commerce Department's Center for AI Standards and Innovation, is resigning after roughly three months, a department spokesperson confirmed. The center, which sits inside NIST and carries responsibility for federal AI evaluation and standards work, has seen repeated turnover in its leadership. The departure is relevant to the field mainly for continuity: standards-and-evaluation bodies shape how government tests and procures AI, and leadership churn slows the standing up of durable measurement and assurance programs.

How it was discussed
  • The Information reported the resignation first and noted continued leadership churn at the standards body; TechCrunch framed the role as a revolving door.
policy standards NIST
#23
Robotic Autonomy 2026-07-20 Hugging Face Blog 6.2 6.4/6.3/6.0

NVIDIA introduced Cosmos 3 Edge, an edge-deployable version of its Cosmos world-foundation-model line aimed at running predictive world-model reasoning directly on robots rather than in the cloud. Pushing world models to the edge targets the latency and connectivity constraints of embodied deployment, letting a robot roll out short-horizon predictions locally for planning and control. It continues the convergence of generative world models with real-time robotic autonomy that has defined much of the embodied-AI stack this year.

Cosmos world model edge robotics
#24
Safety, Policy & Regulation 2026-07-20 MIT Technology Review 6.2 6.1/6.4/6.1

MIT Technology Review reports on research finding that AI resume screeners are more prone than human reviewers to forming systematic biases, raising fresh questions as automated screening spreads across hiring pipelines. The work highlights how models trained on historical hiring signals can amplify rather than neutralize the patterns they inherit, and it lands as a practical counterweight to claims that automated screening is inherently more objective than human judgment.

bias hiring fairness
#25
AI for Science 2026-07-20 arXiv cs.AIarXiv cs.CLarXiv — Evals & BenchmarksarXiv — Generative Media 6.2 6.3/6.3/6.0

Structure-based drug design uses a protein target's 3D structure, often with additional spatial constraints, to generate candidate binders, a space diffusion models have dominated. This work benchmarks language models on the same task under sparse spatial constraints, probing whether sequence-native models can reason about 3D binding geometry when structural information is limited. The benchmark sharpens the comparison between diffusion and language-model approaches to generative molecular design.

drug design SBDD benchmark
#26
Post-Training 2026-07-20 arXiv cs.AIarXiv — Evals & BenchmarksarXiv — Post-trainingarXiv — Reinforcement Learning 6.2 6.3/6.3/6.0

Models trained with the compressed scalar rewards of RLHF often struggle to navigate value conflicts, cases where legitimate objectives pull in opposite directions. This paper takes a geometric view of how chain-of-thought shapes the resolution of such conflicts and proposes a stabilization approach that reduces the erratic behavior scalar rewards induce when tradeoffs are involved. The analysis links reasoning-trace structure to more consistent conflict handling.

RLHF value conflict CoT
#27
Reinforcement Learning 2026-07-20 arXiv cs.AIarXiv — Post-trainingarXiv — Reinforcement Learning 6.2 6.3/6.2/6.0

PPO and the GRPO baseline use clipped surrogate objectives whose favorable-direction saturation creates an abrupt kink in the objective's derivative. OR Else introduces Output Reset, a smooth one-sided saturation that replaces the hard clip with a differentiable trust region, removing the derivative discontinuity that can destabilize policy updates. The result is a cleaner optimization surface for the RL objectives now standard in LLM post-training.

PPO GRPO policy optimization
#28
Efficiency 2026-07-20 arXiv cs.CLarXiv — EfficiencyarXiv — Evals & Benchmarks 6.2 6.3/6.2/6.0

Long-context inference for retrieval-augmented and multi-document workloads is dominated by key-value cache cost. C2KV makes cached KV both compressed and composable, so cache segments from previously processed chunks can be reused and recombined across requests rather than recomputed. By turning the KV cache into reusable building blocks, the method cuts redundant prefill for workloads that repeatedly touch overlapping context.

KV cache long context inference
#29
Interpretability 2026-07-20 arXiv cs.AIarXiv cs.CLarXiv — Mechanistic Interpretability 6.2 6.3/6.2/6.0

Modern LLMs flip correct answers under trivial prompt perturbations: a casual hint, a mislabeled few-shot example, or a fabricated prior assistant turn. This paper probes how alignment tuning changes the internal representations of sycophancy and related failure modes, tracing where the susceptibility lives and how post-training moves it. The mechanistic account connects observable sycophantic behavior to representational structure that alignment does and does not fix.

interpretability sycophancy alignment
#30
Agents & Tool Use 2026-07-20 arXiv — Agents / Tool UsearXiv cs.AIarXiv — Evals & BenchmarksarXiv — Reinforcement Learning 6.1 6.2/6.1/6.0

Video misinformation detection is usually posed as holistic understanding, processing an entire clip and its context in one pass. This work argues real cases often hinge on a few decisive fragments and proposes an agentic evidence-seeking approach that actively locates the sparse spans that settle a claim. Focusing verification on retrieved evidence rather than the whole video improves both accuracy and interpretability of the judgment.

misinformation video agents
#31
Safety, Policy & Regulation 2026-07-20 arXiv — Agents / Tool UsearXiv cs.AIarXiv — Evals & Benchmarks 6.1 6.2/6.1/6.0

Most agent-security benchmarks pit defenders against fixed attack pools collected before evaluation, which understates adaptive threats. Adaptive Adversaries introduces a multi-turn, multi-model benchmark where attackers adjust across a dialogue, exposing agents to evolving prompt-injection and manipulation rather than static attacks. Evaluating defenses against adaptive adversaries yields a more realistic picture of how tool-using agents hold up under sustained pressure.

agent security prompt injection benchmark
#32
Efficiency 2026-07-20 arXiv cs.AIarXiv — EfficiencyarXiv — Evals & Benchmarks 6.1 6.2/6.1/5.9

Can models with substantially different parameter spaces be merged by direct weighted averaging, with no training or semantic alignment? Existing heterogeneous-fusion methods lean on distillation or adapters; this work revisits plain weighted model averaging as a training-free baseline and finds it more competitive than assumed for combining differently-structured models. The result is a cheap merging option worth trying before reaching for heavier fusion machinery.

model merging efficiency
#33
Agents & Tool Use 2026-07-20 arXiv — Agents / Tool UsearXiv cs.AIarXiv cs.RO 6.1 6.2/6.1/5.9

Natural-language control is an appealing interface for UAVs, but self-hosted computer-use agents are built for interactive desktop pacing, not the tight real-time loop flight demands. RT-SHCUA restructures a self-hosted computer-use agent to meet UAV timing constraints, closing the structural mismatch between click-and-wait agent design and continuous vehicle control. It is an early bridge between the computer-use-agent line and embodied real-time autonomy.

computer use UAV real-time
#34
Evaluations & Benchmarks 2026-07-20 arXiv cs.AIarXiv cs.CLarXiv — Evals & Benchmarks 6.0 6.1/6.0/6.0

WorldCupArena evaluates models and deep-research systems on football-match prediction, where a system must integrate changing pre-kickoff information and commit to a clear prediction before the outcome is known. By requiring commitment ahead of ground truth, the benchmark resists the contamination and hindsight problems that plague static evaluations, and it tests genuine forecasting under uncertainty rather than retrieval of settled facts.

benchmark forecasting deep research
#35
Robotic Autonomy 2026-07-20 arXiv cs.ROarXiv — Generative MediaarXiv — Reinforcement Learning 6.0 6.1/6.1/5.8

Rare-failure discovery for autonomous systems has mostly been demonstrated on simple academic driving stacks, leaving open whether it transfers to robust production planners. This work applies importance sampling with PCA-based dimensionality reduction to surface rare failures in more realistic commercial autonomy planners, testing generalization beyond toy stacks. It is a step toward safety-case tooling that works against the planners actually deployed on the road.

autonomous vehicles safety testing
#36
Generative Media 2026-07-20 Hugging Face Daily PapersAK (@_akhaliq) Daily Papers 6.0 6.1/6.0/5.8

The standard diffusion inference framework treats generation as numerical integration and the model as an exact estimator, ignoring the statistical uncertainty in its own predictions. DiFA introduces an inference-time forward-process alignment that accounts for that uncertainty during sampling, adjusting the reverse trajectory rather than assuming the model is exact. The approach improves sample quality without retraining, working purely at inference.

diffusion sampling alignment
#37
Post-Training 2026-07-20 Hugging Face Daily PapersAK (@_akhaliq) Daily Papers 6.0 6.1/6.0/5.8

Token-Level Off-Policy Labeling reframes post-training as a token-level correctness-prediction task rather than sequence-level reward maximization. The intuition is that training a model to distinguish good from bad continuations token by token yields a denser, more faithful signal that holds up better under distribution shift than coarse sequence rewards. The method targets faithful generation where the deployment distribution diverges from training.

post-training off-policy faithfulness
#38
Robotics 2026-07-20 arXiv — Agents / Tool UsearXiv cs.ROarXiv — Evals & Benchmarks 6.0 6.1/6.0/5.8

Long-horizon robot tasks demand diverse skills no single policy provides reliably, but combining heterogeneous policies means reasoning over uncertain, overlapping capability boundaries. RoboHarness uses a memory-driven orchestrator to route sub-tasks among specialized policies, tracking which policy has succeeded where and adapting the hand-offs. It offers a structured way to compose a capable long-horizon system from complementary but imperfect components.

robotics orchestration long-horizon
#39
Multimodal 2026-07-20 Hugging Face Daily PapersAK (@_akhaliq) Daily Papers 6.0 6.2/6.0/5.9

HOMIE addresses subject-driven video generation where both a person and the objects they interact with must stay consistent, a case existing personalization methods handle poorly when the interaction, not just the subject, carries the identity. Using multimodal guidance to jointly preserve human and object identity through interaction, it improves fidelity on human-object-centric video personalization over subject-only approaches. The work targets the harder end of controllable video generation.

video generation personalization multimodal
#40
Audio & Speech 2026-07-20 arXiv cs.CLarXiv — Evals & Benchmarks 5.9 6.0/5.9/5.7

As large audio-language models advance, Spanish speech understanding under realistic, noisy acoustic conditions has received little dedicated evaluation. ESCUCHA introduces a benchmark spanning heterogeneous recording environments to stress-test audio-language models on Spanish beyond clean-studio conditions. It fills a gap in multilingual audio evaluation and gives a reproducible target for robustness work outside English.

speech Spanish benchmark audio
Items
40
Multi-source
29
Long-form (≥7.5)
4
Sources OK / attempted
118 / 119
Top category
Safety, Policy & Regulation
6 items