← Archive / All Digests
A wolf in round glasses reading a book, wrapped in a golden ribbon, in a sunlit forest.

Wolf Digest — Friday, July 10, 2026

Coverage window: 2026-07-09 03:39 ET2026-07-10 03:29 ET
Press play to listen
Friday, July 10, 2026
10m 4s · top-4 narrated briefing
#1 · Frontier LLMs
Meta Superintelligence Labs ships Muse Spark 1.1, a multimodal agentic model, and starts selling API access
Meta opens a paid Model API with Muse Spark 1.1, its multimodal agentic model with a 1M-token context.
8.0 · 3 srcs
#2 · Agents & Tool Use
OpenAI unveils ChatGPT Work, an agentic desktop 'superapp' aimed at Claude Cowork, and sunsets its Atlas browser
OpenAI's ChatGPT Work agent targets Claude Cowork as it retires the Atlas browser and consolidates agentic features.
7.7 · 3 srcs
#3 · Safety, Policy & Regulation
Anthropic opens a 'hard questions' initiative, adds Bernanke to its trust, and ships a Claude usage dashboard
Anthropic launches a public 'hard questions' initiative, adds Bernanke to its trust, and ships a Claude-usage dashboard.
7.6 · 3 srcs
6.5
#1
Frontier LLMs 2026-07-09 Meta AI BlogThe Information — AITechCrunch — AI 8.0 8.2/7.6/8.2

Meta Superintelligence Labs released Muse Spark 1.1, a multimodal reasoning model built for agentic tasks and a substantial upgrade over the original Muse Spark, alongside a public preview of a new, OpenAI-compatible Meta Model API. The API launch is the notable strategic turn: it is the first time Meta has sold access to one of its models, a reversal after years of positioning around open weights. The model is available now in a Thinking mode inside the Meta AI app and on meta.ai.

On capabilities, Muse Spark 1.1 actively manages a one-million-token context window, remembering earlier actions, retrieving information from much earlier in a task, and compacting context to preserve the steps that matter for later work. It zero-shot generalizes to new native tools, Model Context Protocol servers, and custom skills, and it is trained to orchestrate multi-agent systems: as a main agent it gathers context, plans, and delegates execution across parallel subagents, and as a subagent it stays within its assignment and escalates when it hits the edge of its remit. For computer use it decides when to automate and when to act directly, writing scripts when automation is faster, clicking when direct interaction is simpler, and emitting batches of actions per step rather than reasoning one click at a time.

Coding improved markedly on large real-world codebases, covering bug diagnosis, enterprise feature work, and large migrations; on Meta Internal Coding Bench the model beats the original Muse Spark and is described as competitive with leading alternatives, with demos run inside the OpenCode harness. Multimodal strengths include visual-to-code generation and an agent that builds a Facebook Marketplace listing from a smartphone video. Meta says it evaluated the model under its Advanced AI Scaling Framework across chemical and biological, cybersecurity, and loss-of-control categories and reports it within safe margins, with strong resistance to jailbreaks and prompt injection.

The business framing is as important as the model. The Information reports Meta plans to spend heavily to subsidize usage of Muse Spark and pull developers away from rivals, arriving a day after SpaceXAI introduced Grok 4.5 and the same week OpenAI shipped its GPT-5.6 family. Early partners quoted include Replit, Cline, Box, and OpenClaw, praising the million-token context, full multimodal support, built-in search with citations, and parallel tool calling in an OpenAI-compatible package.

How it was discussed
  • Meta's blog leads with the agentic and computer-use gains and the new Meta Model API as a developer on-ramp.
  • The Information frames it as a paid-API pivot away from open source, with Meta subsidizing usage to buy market share.
  • TechCrunch positions it in the crowded AI coding battle, emphasizing bug-fixing and large code migrations.
frontier_llm agents ai_coding
#2
Agents & Tool Use 2026-07-09 The Information — AIOpenAI ResearchTechCrunch — AI 7.7 7.6/7.5/8.0

OpenAI introduced ChatGPT Work, an agent that can take action across a user's apps and files, stay with a project for hours, and turn a goal into finished output. In OpenAI's framing it taps corporate data to automate the creation of spreadsheets and presentations and can handle more involved tasks like updating financial forecasts. The Information describes the release as a direct competitor to Anthropic's Claude Cowork and a step toward a desktop 'superapp' that consolidates OpenAI's business-facing agentic features in one place, part of a broader effort to win enterprise customers.

At the same time, OpenAI is shutting down Atlas, its dedicated AI browser, less than a year after launch. Rather than abandoning agentic browsing, OpenAI is folding those capabilities into its desktop app and a Chrome extension, a signal that the company sees the persistent agent surface, not a standalone browser, as the place where autonomous web actions belong. The consolidation lines up with the ChatGPT Work push: one agent that reaches across documents, apps, and the web.

The competitive backdrop sharpened the same day. The Information separately reported that Cursor is building a general-purpose AI agent to rival Claude Cowork, diversifying beyond its coding roots; that work began after Cursor started leasing compute from SpaceX's AI unit in April and comes ahead of SpaceX's planned sixty-billion-dollar acquisition of Cursor, following the two companies' joint Grok 4.5 launch. Taken together, the day's releases show the frontier labs and their close partners converging on the same product shape at once: a general agent that lives on the desktop, holds a task for hours, and reaches into a company's real data and tools.

How it was discussed
  • The Information casts ChatGPT Work as a Claude Cowork competitor and a consolidating enterprise 'superapp.'
  • OpenAI's own framing stresses hours-long autonomy across apps and files, turning a goal into finished work.
  • TechCrunch notes the Atlas browser shutdown, with agentic browsing moving into the desktop app and a Chrome extension.
agents ai_coding
#3
Safety, Policy & Regulation 2026-07-09 Anthropic NewsThe Information — AITechCrunch — AI 7.6 7.2/7.8/7.8

Anthropic, a public-benefit corporation, opened a 'hard questions' initiative that invites the public to submit their hardest questions about AI, spanning jobs, society, families, science, medicine, and safety, and commits the company to publicly track and report the specific actions it takes to address them, including where it falls short of its stated goals. The company says it has already laid groundwork: an Anthropic Public Record survey that in its first round asked fifty-two thousand Americans for their biggest hopes and concerns, a survey of eighty-one thousand Claude users across a hundred and fifty-nine countries and seventy languages through its Anthropic Interviewer, dozens of in-person focus groups, and analysis of anonymized real-world Claude usage. It also created an Anthropic Institute to study AI's societal challenges, with the Long-Term Benefit Trust providing oversight of the mission.

The governance side moved the same day. Anthropic's Long-Term Benefit Trust, the body charged with appointing board members, named former Federal Reserve chair Ben Bernanke as its fourth member, filling a key seat; the trust still has one of its five seats open. The appointment adds an economist with central-banking experience to the group that exercises impartial oversight of Anthropic's public-benefit obligations.

Anthropic also launched a Reflect dashboard that lets users visualize how they use Claude and judge whether that time aligns with their goals; TechCrunch observed that the feature, while framed as a wellbeing tool, also quietly underscores how much of a user's daily work now depends on the assistant. Separately, Anthropic reiterated that access to Claude models and its coding tools is not permitted in China, responding to the Chinese government's recent warning about a claimed security backdoor in Claude Code, and pushed back on Beijing's characterization. The threads share a posture: a lab leaning into public-benefit framing, external oversight, and transparency about usage even as it defends its deployment boundaries.

How it was discussed
  • Anthropic's own announcement centers public input and a commitment to report progress and shortfalls.
  • The Information frames the Bernanke appointment as filling a key governance seat on the Long-Term Benefit Trust.
  • TechCrunch reads the new Reflect dashboard as also reinforcing user dependence on Claude.
safety_policy industry
#4
AI for Science 2026-07-09 Google AI Blog 7.5 7.6/7.6/7.3

Google presented SensorFM, a foundation model for wearable health pretrained on more than one trillion minutes of sensor data from five million people using missing-aware masked reconstruction over multimodal signals. The central claim is that co-scaling model size and data yields a general-purpose representation of human physiology that transfers broadly, supports label-efficient adaptation and data infilling, and can serve as a grounding tool for a personal health agent, with performance gains that show no sign of saturating as compute grows.

To probe generality, the team evaluated SensorFM across thirty-five discriminative health tasks drawn from three independent, institutional-review-board-approved prospective studies totaling nearly fourteen thousand participants, spanning cardiovascular health, metabolic risk, mental health, sleep, demographics, and lifestyle. Keeping the encoder frozen and training only a lightweight linear head, SensorFM embeddings beat a feature-engineered supervised baseline on thirty-four of thirty-five tasks with no task-specific architecture. Adding demographic features like age and sex gave only a modest boost that shrank as the model scaled, evidence that larger models implicitly capture physiologically relevant traits during pretraining. The gains were largest on hard-to-measure conditions such as depression and anxiety, which leave faint, person-specific traces in sensor data, and the model reached strong accuracy with only a small fraction of labeled examples, an important property where high-quality labels are scarce.

Google also automated the usually tedious step of turning embeddings into predictors. It built an agentic 'classroom' of collaborating and competing language-model agents that iteratively generate, test, and refine executable code to construct prediction heads on the SensorFM embeddings, exploring more than thirty thousand candidate solutions. The agent-designed heads beat a simple linear probe on sixteen of twenty classification tasks and twelve of fifteen regression tasks, with solution quality improving monotonically over the search and scaling with the capability of the underlying model, so that more capable Gemini versions produced better heads while agent collaboration helped weaker models close the gap. The company frames SensorFM as a shift away from many bespoke single-outcome models toward one generalist representation of physiology; the work is research rather than a cleared medical device.

ai_science frontier_llm research
#5
Industry 2026-07-09 OpenAI ResearchTechCrunch — AI 7.1 7.0/7.0/7.2

OpenAI said its newly launched GPT-5.6 family is now the preferred model powering Microsoft 365 Copilot across Word, Excel, PowerPoint, Chat, and Copilot's agentic surfaces. The announcement lands amid continued 'breakup chatter' over the two companies' relationship, and functions as a public signal that OpenAI's models will keep driving Microsoft's flagship productivity suite even as both firms diversify their partners. It also gives GPT-5.6 an immediate, enormous distribution channel days after the family reached general availability.

How it was discussed
  • OpenAI's post emphasizes capability gains across Office apps; TechCrunch frames the same news against ongoing OpenAI-Microsoft 'breakup' speculation.
frontier_llm industry
#6
Industry 2026-07-09 AI Snake Oil (Narayanan & Kapoor) 6.9 6.9/7.2/6.6

In a widely shared essay, Arvind Narayanan and Akash Kapur argue that the debate over whether AI is a bubble misses the more consequential question of who captures value long term. Their thesis: the labs are not confined to being model providers and are aggressively migrating up the stack into applications, agents, and enterprise platforms. That move likely lets them escape the commodity trap that pure model-serving implies, but it raises new concerns around customer lock-in and reduced competition. The piece is a useful analytic frame for the same day's product news, in which Meta opened a paid API, OpenAI pushed an agentic 'superapp,' and Cursor began building a general agent.

industry safety_policy
#7
Government & Defense 2026-07-09 DefenseScoop 6.8 6.9/6.8/6.6

Accenture Federal Services was selected for a five-year task order worth up to eight hundred twenty-one million dollars to supply core integration support for the Pentagon's War Data Platform, the Chief Digital and AI Office program that grew out of the Advana enterprise data and analytics effort. According to government spending records and market-intelligence trackers cited by DefenseScoop, Accenture beat four other commercial bidders. The award anchors the data-integration backbone the department relies on for analytics and AI across its enterprise, and signals continued consolidation of that work under a single prime after the Advana-to-War-Data-Platform rebrand.

gov_defense industry
#8
Agents & Tool Use 2026-07-09 The Information — AI 6.8 6.9/6.7/6.9

Cursor is developing a general-purpose AI agent meant to rival Anthropic's Claude Cowork, part of a push to diversify beyond coding-focused tools, according to The Information. Work reportedly began after Cursor started leasing compute from SpaceX's AI unit in April and comes ahead of SpaceX's planned sixty-billion-dollar acquisition of Cursor; the new agent could strengthen the SpaceXAI enterprise business once the deal closes, building on the two companies' joint Grok 4.5 launch. It is another data point in a broad convergence toward general desktop agents across the frontier ecosystem.

agents industry
#9
Safety, Policy & Regulation 2026-07-09 The Information — AI 6.7 6.7/6.9/6.4

Anthropic reiterated that access to its Claude models and coding tools is not permitted in China, responding to a Chinese government warning about a claimed security backdoor in Claude Code. The company said what Beijing described as a hidden backdoor is not one, and reaffirmed its existing policy restricting Chinese access. The exchange highlights the widening entanglement of frontier-model deployment with export-control and national-security politics between the two governments.

safety_policy industry
#10
Industry 2026-07-09 The Information — AI 6.7 6.6/6.9/6.5

At least twenty-two professors and researchers have taken leave or stepped back from roles at Stanford, Berkeley, Harvard, Virginia, and the University of Southern California this year to join OpenAI, Anthropic, Google, or Meta, according to a review by The Information, with the true figure likely higher. Remaining faculty say the departures threaten open-source AI development in the West and hollow out academic labs that train the next generation of researchers, as frontier labs outbid universities for scarce senior talent.

industry
#11
Evaluations & Benchmarks 2026-07-02 AK (@_akhaliq) Daily PapersHugging Face Daily Papers 6.7 6.8/6.6/6.6

Rather than propose another benchmark, Video-Oasis is a diagnostic suite that audits existing video-understanding benchmarks for shortcuts. Its central finding is that fifty-five percent of samples across current benchmarks are solvable without visual input or temporal context, meaning reported Video-LLM scores conflate genuine perception with linguistic priors and knowledge. After filtering these shortcut samples, the remaining set gives a cleaner read on temporal visual reasoning, and the tool offers a reusable way to vet future benchmarks before trusting their numbers.

cs.CV evals
#12
Efficiency 2026-07-09 The Information — AI 6.5 6.7/6.5/6.4

Stealth-emerged startup PrismML says it compressed Alibaba's open-source Qwen 3.6, a twenty-seven-billion-parameter model, to run on an iPhone 17 Pro, which it claims is the largest model yet run on a phone. On-device inference of a model this size would cut cloud costs and improve privacy, the pitch Apple itself has been chasing. The specific compression method was not detailed; the claim, if it holds up under independent testing, pushes the ceiling on local large-model inference for consumer hardware.

efficiency industry
#13
Industry 2026-07-09 The Information — AITechCrunch — AI 6.5 6.4/6.6/6.6

Fidji Simo said she is stepping down as OpenAI's CEO of AGI deployment as she manages a chronic illness, and will transition to a part-time advisory role. Her departure reopens a senior leadership vacancy at a company already reshuffling its executive ranks, and removes one of OpenAI's most prominent operators from a day-to-day post as it scales its consumer and enterprise businesses.

How it was discussed
  • The Information broke the transition and its effect on OpenAI's leadership bench; TechCrunch confirmed the move to a part-time advisory role.
industry
#14
Industry 2026-07-09 TechCrunch — AI 6.5 6.4/6.4/6.6

Ollama, the open-source tool that makes it easy to run models locally, raised sixty-five million dollars in a Benchmark-backed round, reporting nearly nine million users and roughly one hundred seventy-six thousand GitHub stars. The funding underscores durable developer demand for running open-weight models on their own machines, a counterweight to the frontier labs' push toward hosted, API-gated agents, and positions Ollama as a default local runtime as more capable open checkpoints ship.

industry efficiency
#15
Government & Defense 2026-07-09 DefenseScoop 6.5 6.6/6.5/6.3

The Defense Department awarded other-transaction agreements worth eighty-six million dollars to nLIGHT Defense and Lockheed Martin Aculight for directed-energy weapons under the Joint Laser Weapon System program, run by the Pentagon's research and engineering directorate, aimed at defeating adversary drone swarms and cruise missiles. Officials cite high-speed engagement, low cost-per-shot, and deep magazines as directed-energy advantages, while acknowledging the department's long struggle to move such systems from research to fielded production.

gov_defense industry
#16
Safety, Policy & Regulation 2026-07-09 TechCrunch — AI 6.5 6.4/6.6/6.4

The New York Times and other news publishers filed a motion for sanctions alleging OpenAI withheld tools and datasets that could identify copyrighted journalism reproduced in ChatGPT outputs, escalating their copyright suit. The dispute turns on discovery over training data and output attribution, and its outcome could shape how courts treat evidence of memorized or reproduced copyrighted text in generative systems.

safety_policy industry
#17
Generative Media 2026-07-03 AK (@_akhaliq) Daily PapersHugging Face Daily Papers 6.5 6.6/6.3/6.7

Vidu S1 is a real-time interactive video generation model that lets users steer content at any moment through voice instructions and supports infinite-length generation without blurring, drift, or visual distortion. Built on the authors' TurboDiffusion and TurboServe stack, it outputs 540p video at up to forty-two frames per second on regular consumer GPUs and accepts custom images of people, anime, or pets with selectable voice tones. The paper reports best-in-class scores across its test metrics while meeting real-time inference constraints, and it drew unusually strong community attention on the daily-papers feed.

cs.CV generative_media
#18
AI for Science 2026-07-09 AK (@_akhaliq) Daily PapersarXiv cs.AI (Artificial Intelligence)Hugging Face Daily Papers 6.5 6.6/6.6/6.4

IG-Bench treats scientific ideas like genomes that inherit mechanisms, repair prior limitations, and recombine earlier work. Each paper or proposal is represented as typed, evidence-grounded Idea Genome objects, and a GenomeDiff aligns them to record inheritance, mutation, loss, external import, and novel insertion under six evolutionary dynamics. The benchmark ships nearly two thousand curated golden lineage traces, and asks whether models can both trace an idea's ancestry and generate new proposals that are properly grounded in that lineage, a more structured probe of scientific reasoning than open-ended idea prompting.

cs.AI evals
#19
Infrastructure 2026-07-09 TechCrunch — AI 6.4 6.5/6.4/6.3

Meta's next-generation in-house AI chips are slated to begin production in September, with the company taking a deliberately modular design approach on the expectation that its workloads will shift as models evolve before the silicon ships. The effort is part of Meta's push to reduce dependence on Nvidia for training and inference, and dovetails with reports the same day that Meta intends to spend heavily to subsidize usage of its new Muse Spark model.

infra industry
#20
AI for Science 2026-07-09 Microsoft Research Blog 6.4 6.5/6.5/6.1

Microsoft Research released Aurora 1.5, an update to its open foundation model for the atmosphere that extends coverage from medium-range weather toward wider Earth-system applications. Building on Aurora's pretrain-then-fine-tune recipe over large volumes of geophysical data, the release broadens the set of environmental variables and downstream tasks the model targets, continuing the trend of a single foundation model replacing many bespoke, physics-only forecasting pipelines while remaining openly available to researchers.

ai_science research
#21
Safety, Policy & Regulation 2026-07-09 TechCrunch — AI 6.4 6.4/6.6/6.1

TechCrunch examines the opaque process behind the government-gated pre-release review that cleared OpenAI's GPT-5.6 family, noting that the exact dialogue between federal authorities and the labs remains undisclosed. The piece raises the accountability question created by informal safety sign-offs: as governments take a hand in approving frontier releases, the criteria, evaluators, and evidence involved are not yet transparent to the public or the research community.

safety_policy industry
#22
Multimodal 2026-07-09 AK (@_akhaliq) Daily PapersarXiv cs.CV (Computer Vision) 6.4 6.5/6.4/6.4

OpenCoF explores a reasoning mode distinct from text chain-of-thought: Chain-of-Frame, in which a video generation model reasons by unrolling temporally connected frames. Because general video generators lack supervision for this, the authors release OpenCoF-17K, a reasoning-video dataset spanning eleven task families, and Wan-CoF, a fine-tuned video model, to test whether diverse temporal supervision improves frame-by-frame reasoning. The framework provides both data and a model for studying whether video generation can serve as a substrate for logical, consequence-aware reasoning rather than only appearance modeling.

cs.CV multimodal
#23
Evaluations & Benchmarks 2026-07-09 AK (@_akhaliq) Daily PapersarXiv cs.CL (Computation & Language)Hugging Face Daily Papers 6.4 6.5/6.4/6.4

UniClawBench is a capability-driven benchmark for proactive agents that operate everyday tools, built to fix two weaknesses in prior evaluations: reliance on sandboxed, single-turn settings and task taxonomies that entangle multiple capabilities. Organized around five foundational agent capabilities in dynamic multi-turn environments, it isolates where agents fail rather than reporting a single blended score, giving a clearer diagnostic picture of proactive-agent competence on realistic tasks.

cs.CL evals
#24
Multimodal 2026-07-09 arXiv cs.CV (Computer Vision)arXiv cs.AI (Artificial Intelligence)arXiv cs.CL (Computation & Language)arXiv cs.LG (Machine Learning) 6.4 6.5/6.3/6.4

Unified multimodal models that understand, generate, and edit images usually replay all past visual and textual inputs into one context window, which explodes visual-token counts and makes cross-turn reference unreliable over long dialogues. This work externalizes visual information into an episodic visual memory and selectively reactivates relevant episodes during reasoning, using a perceptual abstraction engine for structured visual abstraction, a cognitive retrieval engine for cross-turn memory, and an executive controller for task inference. The design targets long-horizon multimodal dialogue without the token blow-up of naive context concatenation.

cs.CV multimodal agents
#25
Interpretability 2026-07-09 arXiv cs.CV (Computer Vision)arXiv cs.AI (Artificial Intelligence)arXiv cs.LG (Machine Learning) 6.4 6.6/6.5/6.1

In vision-language models, vanilla sparse autoencoders often learn concepts that fragment in the visual modality, covering disjoint image regions even when the language concept is coherent. The proposed Structured Sparse AutoEncoder enforces concept consistency both semantically and spatially: it groups image patches by transformer attention similarity and spatial proximity and adds a structured sparsity regularizer during training. The result is SAE features that align more cleanly across modalities, a step toward mechanistic interpretability that holds up when a single concept must be recovered from both pixels and text.

cs.CV interpretability
#26
Robotic Autonomy 2026-07-09 Shield AI 6.3 6.5/6.4/6.1

Shield AI and Destinus integrated Hivemind mission-autonomy software onto the Destinus Hornet platform, paving the way for its use on the Destinus Ruta deep-strike cruise missile. The pitch, drawn from recent Ukraine and Middle East combat, is autonomy that operates through jamming, degraded links, and shifting targets while coordinating with intelligence, surveillance, and relay assets in real time. Shield AI says Hivemind now runs on more than thirty platforms, positioning it as a widely portable edge-autonomy stack for contested environments.

robotic_autonomy gov_defense
#27
Audio & Speech 2026-07-09 TechCrunch — AI 6.3 6.3/6.2/6.3

Paris-based AI voice startup Gradium raised a hundred-million-dollar seed round backed by Nvidia, and plans to open a Bay Area office to compete for talent. The unusually large seed reflects continued investor appetite for speech and audio foundation models, a segment drawing heavy funding as real-time voice agents become a core product surface for the major labs.

audio industry
#28
Efficiency 2026-07-09 AK (@_akhaliq) Daily PapersarXiv cs.CV (Computer Vision) 6.3 6.4/6.3/6.2

Few-step autoregressive video diffusion models generate long videos with low latency but accumulate error and lose motion dynamics over long rollouts. OPSD-V is an on-policy self-distillation scheme that has the student follow the exact inference-time rollout, generating each chunk conditioned on its own previously generated KV cache, while a teacher provides dense trajectory-level supervision using real long-video context. This preserves the fast few-step inference path while reducing long-horizon degradation, targeting the specific failure mode that makes streaming video generators drift.

cs.CV efficiency generative_media
#29
Robotics 2026-07-09 arXiv cs.RO (Robotics)arXiv — Agents / Tool UsearXiv — Robotic Autonomy / Embodied AI 6.3 6.4/6.3/6.1

DexVerse is a large, modular benchmark for dexterous manipulation spanning a hundred tasks, including grasping and relocation, articulated-object interaction, functional tool use, bimanual coordination, non-prehensile control, and contact-rich behaviors. It is designed to test cross-task and cross-embodiment generalization with controllable visual variation, addressing the narrow task and embodiment coverage of prior benchmarks and giving manipulation-policy researchers a systematic way to measure how well skills transfer across robots and conditions.

cs.RO robotics evals
#30
Audio & Speech 2026-07-09 arXiv cs.CL (Computation & Language)arXiv cs.AI (Artificial Intelligence)arXiv — Reinforcement Learning 6.3 6.4/6.3/6.2

Adapting LLM-based speech recognition to regulated domains like banking is bottlenecked by privacy, pushing teams toward synthetic text-to-speech that stays acoustically mismatched with real audio. This work shows Group Relative Policy Optimization, a critic-free method rewarding low word-error-rate hypotheses, extracts far more from the same synthetic speech than supervised fine-tuning, cutting word error rate by forty percent relative to fine-tuning, from 36.71 to 22.09 percent, and reaching forty-five percent when fine-tuning is followed by GRPO. It is a concrete case where reinforcement learning beats supervised adaptation on distribution-shifted synthetic data.

cs.CL audio rl
#31
Robotics 2026-07-09 C4ISRNET 6.2 6.3/6.3/6.0

Ukrainian unmanned ground vehicle maker Trinity Robotics plans to roughly double production to about twenty-two hundred units this year as the Ukrainian army leans harder on ground robots, and is in talks with a French manufacturer and other European partners to build abroad. The ramp illustrates how battlefield demand is scaling combat-robot manufacturing and pulling Ukrainian autonomy suppliers into cross-border defense-industrial ventures.

robotics gov_defense
#32
Infrastructure 2026-07-09 The Information — AI 6.2 6.3/6.3/5.9

Lancium, the Blackstone-backed power-infrastructure developer behind OpenAI and Oracle's Texas data-center campus, is in talks to sell a minority stake, with interest from tech companies including Nvidia, according to The Information. Its portfolio of gigawatt-scale Texas campuses sits at the center of the compute-and-power buildout underpinning frontier training, and the deal talks underscore how energy siting has become a strategic asset in the AI infrastructure race.

infra industry
#33
Multimodal 2026-07-09 arXiv cs.CV (Computer Vision)arXiv — Evals & BenchmarksarXiv — Generative Media / Diffusion 6.2 6.3/6.2/6.1

Unified multimodal models that interleave textual and visual reasoning typically render each intermediate visual state as a full image, wasting tokens and diluting supervision on the small, reasoning-critical changes between steps. DeltaV instead predicts compact update tokens conditioned on prior visual states, capturing only what changes across reasoning steps and sizing the token budget to the magnitude of visual change. The approach reduces visual-token redundancy while sharpening supervision on the transitions that actually carry the reasoning.

cs.CV multimodal
#34
Efficiency 2026-07-09 arXiv cs.LG (Machine Learning)arXiv cs.AI (Artificial Intelligence)arXiv — Evals & Benchmarks 6.2 6.3/6.3/6.0

Standard speculative decoding is lossless, exactly preserving the target model's sampling distribution, but recent work argues relaxing that guarantee can buy extra speed or tunable trade-offs. This paper unifies training-free relaxed variants in one framework and benchmarks them on contemporary settings, distilling practitioner takeaways, chief among them that relaxation can cost meaningful capability and must be applied carefully. It is a useful map of when loosening the exactness guarantee is worth it and when it quietly degrades quality.

cs.LG efficiency
#35
Interpretability 2026-07-09 arXiv cs.CL (Computation & Language)arXiv — Evals & BenchmarksarXiv — Mechanistic Interpretability 6.2 6.3/6.3/6.0

Because dictionary learning is non-convex, independently trained networks learn misaligned feature spaces, so apparently identical concepts can differ purely by random seed, a core obstacle to claims of feature universality in interpretability. This method computes an orthogonal Procrustes rotation between seeds' activation spaces before jointly training a Top-K sparse autoencoder end to end, with a dead-feature revival loss. Evaluated on five seed pairs across ten BERT models, it extracts cross-seed universal features, giving evidence for which learned concepts are genuinely shared rather than initialization artifacts.

cs.CL interpretability
#36
Efficiency 2026-07-09 arXiv cs.LG (Machine Learning)arXiv — Efficiency (Quantization, MoE, Inference) 6.2 6.4/6.3/5.9

Scalar and group-wise quantization run out of representational capacity near two bits per weight, while vector quantization adds explicit codebooks, index lookups, and storage overhead. BiSCo-LLM is a codebook-free binary spherical coding framework for extreme low-bit weight compression that keeps a richer block-level representation without lookup tables, targeting the memory-capacity, weight-bandwidth, and checkpoint-storage pressures of large-model deployment at aggressive bit budgets.

cs.LG efficiency
#37
AI for Science 2026-07-09 arXiv cs.AI (Artificial Intelligence)arXiv cs.LG (Machine Learning)arXiv — Reinforcement Learning 6.2 6.3/6.3/6.0

Most generative drug-design models condition only on a target or generic molecular properties, ignoring how disease context shapes target behavior and outcomes. DrugGen-2 designs small molecules conditioned on both disease ontology and target protein sequence, fine-tuning a GPT-2 backbone on approved drugs linked to their diseases and targets, then applying reinforcement learning via Group Relative Policy Optimization with reward functions for chemical validity and other design criteria. The disease-aware conditioning is the contribution, aiming to generate candidates that respect therapeutic context rather than target binding alone.

q-bio.QM ai_science rl
#38
Recurrent & Linear Attention 2026-07-08 AK (@_akhaliq) Daily PapersHugging Face Daily Papers 6.2 6.4/6.3/5.9

Linear-attention models keep a fixed state size and per-token compute but trail softmax transformers on long-context recall, and simply enlarging the state raises FLOPs. Sparse Delta Memory extends the Gated DeltaNet architecture by replacing the dense key-value outer product with sparse reads and writes to a large explicit memory, scaling hidden-state capacity by orders of magnitude while holding compute roughly fixed. Under an isoFLOP comparison it targets the recall gap that has kept linear-attention architectures behind transformers on long sequences.

cs.LG recurrent ssm
#39
Interpretability 2026-07-09 arXiv cs.AI (Artificial Intelligence)arXiv cs.CL (Computation & Language) 6.2 6.3/6.3/6.0

Finetuning can make an LLM memorize a new fact quickly yet still fail to use it in downstream reasoning, a failure the authors formalize as the Knowing-Using Gap, marked by both an accuracy gap and a temporal lag between memorization and generalization. Using a self-patching intervention that relocates internal representations to find where generalization breaks, they present evidence for a knowledge-circuit misalignment: memorized representations sit in the wrong place for the circuits that would use them. The work is a step toward mechanistically explaining, and eventually fixing, brittle knowledge injection.

cs.AI interpretability post_training
#40
Evaluations & Benchmarks 2026-07-09 arXiv cs.CL (Computation & Language)arXiv — Evals & BenchmarksarXiv — Reinforcement Learning 6.2 6.2/6.3/6.0

As reinforcement learning increasingly uses an LLM judge as the reward model, the judge's capability and bias matter directly. This paper studies that calibration for citation quality in deep-research systems, where each written claim must be supported by a cited source judged on relevance and factual support. Scoring eight off-the-shelf judges from three model families against gold labels over more than twelve hundred human-verified rubric decisions on an adversarial long-form benchmark, it quantifies how capable a citation verifier must be before its signal can be trusted for training.

cs.CL evals rl
#41
Government & Defense 2026-07-09 Defense Innovation Unit (DIU) 6.1 6.1/6.2/5.9

The Defense Innovation Unit, with U.S. Indo-Pacific Command, U.S. Northern Command, and the Defense Logistics Agency, awarded two prototype contracts under a Joint Sustainment Decision Tool project to modernize decision-making for the Department of War's Joint Logistics Enterprise. The goal is to move military logistics from reactive processes to a predictive, proactive posture, applying analytics and AI to sustain supply chains at speed in contested environments during large-scale combat operations.

gov_defense industry
#42
Agents & Tool Use 2026-07-09 arXiv cs.CL (Computation & Language)arXiv cs.AI (Artificial Intelligence)arXiv — Agents / Tool Use 6.1 6.2/6.2/6.0

A single ReAct-style search agent is limited by one long trajectory and finite context, while flat multi-agent systems parallelize coverage but handle recursive depth and adaptive collaboration poorly. WebSwarm is a progressive recursive-delegation framework that jointly builds task decomposition, recursive expansion, and agent collaboration at inference time, dynamically instantiating agents to pursue both breadth and depth with evidence-grounded expansion. It targets the depth-versus-coverage bottleneck in agentic web research.

cs.CL agents
#43
Reinforcement Learning 2026-07-08 AK (@_akhaliq) Daily PapersHugging Face Daily Papers 6.1 6.3/6.2/5.8

Reinforcement learning for LLM reasoning relies on importance sampling, which trades off catastrophic instability against clipping that throttles exploration. Formalizing a notion of Probability Capacity, the authors show conservative clipping structurally stifles exploration by prematurely truncating the update budget for correct but low-confidence reasoning paths. Their Unbounded Positive asymmetric optimization loosens that constraint to preserve exploration without the instability, aiming for more sample-efficient RL training of reasoning models.

cs.LG rl post_training
#44
Efficiency 2026-07-08 AK (@_akhaliq) Daily PapersHugging Face Daily Papers 6.1 6.2/6.2/5.9

Long-context deployment often pushes inputs an order of magnitude past the pretraining window, making zero-shot context extension the default for open-weight checkpoints, but fixed rescaling factors either hurt short-context fidelity or break at long range. Jet-Long is a tuning-free method that pairs a local, RoPE-faithful window with a long-range window whose rescaling factor adapts to the current sequence length, recovering short-context behavior while remaining stable at very long contexts.

cs.CL efficiency frontier_llm
#45
AI for Science 2026-07-04 AK (@_akhaliq) Daily PapersHugging Face Daily Papers 6.1 6.2/6.2/5.8

TESSERA v2 reports the largest controlled scaling study to date for Earth-observation foundation models: three hundred ninety-five training runs on more than a thousand GH200 superchips within a fixed pixel-wise Barlow Twins family, each evaluated on fifteen downstream tasks. A striking finding is that pretraining loss barely predicts downstream performance, with absolute Pearson correlation under 0.2, so selecting models by loss wastes compute; the study also finds encoder and data should grow together while the projector stays fixed, giving a concrete compute-allocation rule for remote-sensing foundation models.

cs.CV ai_science
Items
45
Multi-source
27
Long-form (≥7.5)
4
Sources OK / attempted
94 / 119
Top category
AI for Science
5 items