← Archive / All Digests
A wolf in round glasses reading a book, wrapped in a golden ribbon, in a sunlit forest.

Wolf Digest — Monday, August 3, 2026

Coverage window: 2026-08-01 03:58 ET2026-08-03 03:02 ET
Press play to listen
Monday, August 3, 2026
13m 31s · top-4 narrated briefing
#1 · Safety, Policy & Regulation
TruffleHog scan of 7.6 petabytes of public Hugging Face datasets finds 221,303 live credentials across 6,003 corpora
Truffle Security cloned and scanned every public dataset on Hugging Face — 7.6 petabytes, 186.9 million unique files, roughly 815,000 dataset repositories, of which about 670,000 finished cleanly — flattening Parquet, Arrow, JSONL, archives and binaries into scannable text and ru…
8.5 · 2 srcs
#2 · Safety, Policy & Regulation
EU AI Act model-layer obligations become enforceable, with transparency, copyright disclosure and content-labelling duties
The European Union's AI Act provisions governing AI models — as distinct from AI applications — became enforceable on 2 August 2026, and the practical effect is that the European Commission now sits as the most prominent AI regulator in the world. The Act itself passed in 2024, b…
8.4 · 3 srcs
#3 · Government & Defense
AWS security leadership describes adversaries poisoning open-source AI libraries with conditional triggers and AI-built maintainer histories
Defense One's science and technology editor reports on a shift that has happened quickly: the same hyperscalers that argued in 2024 that closed models were safer are now underwriting open weights, and attackers have moved with them. A 24 July 2026 statement from Nvidia's chief ex…
8.0 · 1 srcs
6.5
#1
Safety, Policy & Regulation 2026-08-02 Hacker News — AI front pageTruffle Security 8.5 8.5/8.5/8.4

Truffle Security cloned and scanned every public dataset on Hugging Face — 7.6 petabytes, 186.9 million unique files, roughly 815,000 dataset repositories, of which about 670,000 finished cleanly — flattening Parquet, Arrow, JSONL, archives and binaries into scannable text and running TruffleHog with live verification enabled. The result is 221,303 unique, verified-live credentials sitting inside 6,003 public training corpora. The previous largest web-scale scan the team had run topped out near 400 terabytes, so this is roughly a nineteen-fold jump in surface area.

The breakdown matters more than the headline. Cloud infrastructure accounted for 13,100 live keys: 8,557 Google Cloud service-account keys across 3,811 projects, including 1,926 Firebase admin keys, one with an explicit Owner role and one Kubernetes cluster-admin; 3,343 AWS keys that passed STS identity checks, 907 of which could list buckets, with 8,676 buckets visible through metadata and a confirmed 51.7 terabytes sitting in buckets that had public access blocked. Hosted databases contributed 8,594 live logins totalling 3.5 terabytes by size metadata — the median MongoDB cluster was a trivial 2.8 megabytes, but 89 MongoDB clusters and five Postgres databases exceeded a gigabyte, the largest MongoDB instance holding 617.7 gigabytes. Email and messaging contributed 14,500 keys, and AI provider accounts another 10,700, including 742 live OpenAI keys and 26 live Anthropic keys. Using published entry-tier spend caps alone, the authors put a floor of roughly $920,000 a year on that exposure, and they are explicit that this is a floor derived from caps, not a measurement of actual balances or unauthorized use.

The supply-chain findings are the sharper edge. Among 349 live GitHub personal access tokens, 223 carried full repository write, 130 could rewrite CI workflows, 112 held admin:org, and 110 could publish packages; 318 Docker Hub tokens carried push rights. One live repository-scoped token belonged to the founder of a widely used Model Context Protocol registry whose connected organization repositories carry more than 178,000 GitHub stars combined. npm and PyPI were checked specifically and returned zero live tokens.

The amplification analysis is what makes this structural rather than anecdotal. Forty-four percent of unique live secrets appear in more than one dataset, and 19,380 appear in ten or more. The Stack and its forks alone carry 51,571 distinct live keys; the Dolma family carries 28,110; and of keys reaching ten or more datasets, 99.3 percent pass through one of these scrape corpora. Two near-identical stack-edu re-uploads share all 19,977 of their keys. Chat logs turn out to be a genuinely new leak path: one Infura key pasted into a ChatGPT conversation was captured by WildChat and copied into 1,131 public datasets across 10,162 file locations, and a live AWS key for a Brazilian lending fintech reached training data through boto3 code pasted into a chatbot, captured by LMSYS-Chat-1M and mirrored roughly eighteen times.

Impact measurement was deliberately metadata-only — database size statistics, Redis memory counters, CloudWatch bucket-size metrics — with no rows read and no objects listed. Coverage is incomplete, closed corpora are invisible by construction, and disclosure at this scale is unsolved: emailing 221,303 owners is not realistic, so names, dataset names, file paths and key material are all being withheld while notification runs through provider partner channels. Hugging Face's CTO contributed native storage-bucket scanning support to TruffleHog for an upcoming release.

security training data supply chain Hugging Face
#2
Safety, Policy & Regulation 2026-08-02 Hacker News — AI front pageEuronewsEngadget 8.4 8.0/9.0/8.2

The European Union's AI Act provisions governing AI models — as distinct from AI applications — became enforceable on 2 August 2026, and the practical effect is that the European Commission now sits as the most prominent AI regulator in the world. The Act itself passed in 2024, but the model-layer obligations were the ones held back, and they were only added at all after the public launch of ChatGPT pushed policymakers to extend a framework originally scoped to applications down into the underlying technology.

The general-purpose obligations attach to any model that lacks a specific purpose but can be adapted to a variety of use cases. There are three: transparency on how the model was built, disclosure of any copyright-protected content used in training, and enough information for downstream users to understand what the model can do. A further layer applies to models designated as frontier, whose developers must identify and mitigate risks to society at large. The Commission endorsed a voluntary Code of Practice last year, drafted by outside experts including Yoshua Bengio, spelling out what compliance looks like in practice; most leading Western labs signed it, with Meta the notable exception. OpenAI's head of policy for Europe, the Middle East and Africa said the company has collaborated closely with the Commission on implementation, including the Codes of Practice.

Running alongside is a labelling duty that took effect the same day. Companies must label generated images, audio and text that were designed to look authentic, attaching a digital watermark. Personal content is out of scope, as are works that are evidently artistic, satirical or fictional. New systems placed on the EU market had to comply on 2 August; systems already on the market get four additional months. Non-compliance can draw fines of up to three percent of total gross revenue. The Union has published its own black-and-white labels for anyone to use, and organisations may design their own instead. For scale on the underlying tooling, Google's SynthID has been used to label more than 100 billion images and roughly 60,000 years' worth of audio.

Enforcement falls to the European AI Office, which was created for exactly this and which plans to supplement limited in-house capacity with a panel of scientists and a pool of specialised AI safety firms. The Act reaches foreign companies commercialising AI technologies in the Union, so the practical scope is considerably wider than European developers. Two framings are competing for the Office's attention: the AI ethics tradition concerned with discrimination, privacy and human oversight, and a risk tradition concerned with weapons uplift, large-scale cyberattacks and loss of control. Two recent incidents are cited as potentially steering priorities toward the latter — Anthropic's Mythos-based model being pulled under United States export control restrictions over its cyber capabilities, and an OpenAI agent hacking into an AI firm during testing.

Several practical frictions are already visible. The Office's resources are described as knowingly limited, and frontier models are a moving target that officials must track without much prior experience or scientific consensus on preventing harm at scale. Industry criticism reported alongside the rollout argues the law shifts spending from engineers to lawyers. The most concrete consumer-facing prediction is that the most advanced models may launch in the Union a few weeks later than elsewhere. No compute threshold or euro fine cap for the model-layer duties is specified in the reporting.

How it was discussed
  • Euronews frames the AI Office's enforcement capacity as the binding constraint, competing with industry for scarce AI talent.
  • Engadget centres the labelling duty, noting the three percent gross-revenue penalty and the four-month grace period for pre-existing systems.
  • CCIA's Boniface de Champris told The Guardian consumers will be surprised to see labels in advertising, film and publishing, where AI is already used at scale.
  • CDT's Laura Lazaro Cabrera cautioned against devoting enforcement resources solely to cyber-offence and loss-of-control risks.
EU AI Act regulation watermarking transparency
#3
Government & Defense 2026-08-02 Defense One 8.0 6.8/7.6/6.5 +1.0 gov_defense

Defense One's science and technology editor reports on a shift that has happened quickly: the same hyperscalers that argued in 2024 that closed models were safer are now underwriting open weights, and attackers have moved with them. A 24 July 2026 statement from Nvidia's chief executive, co-signed by Amazon, Meta, Google and Microsoft, defines open-weight models as models anyone can download, inspect, modify and run on their own infrastructure, and calls them an important part of the foundation. The contrast with 2024 is sharp — Microsoft's chief executive then described closed source as critical to safety precisely because it allowed deep end-to-end red-teaming and alignment evaluation before public exposure, at a point when Microsoft had put thirteen billion dollars into OpenAI. Amazon Web Services completed a fifty-billion-dollar investment in OpenAI this week.

Four drivers are offered for the repositioning: research showing cheap open-weight models steadily closing on closed-model performance; rising American public pessimism about AI; the Pentagon wanting AI it can control and run without large, targetable data centres; and the hyperscalers' own move from selling models to selling the tools, space and security that open developers need. The scale numbers from earnings calls make the second business model concrete — more than ten models on Amazon Bedrock, and a Microsoft catalogue of over eleven thousand models spanning OpenAI, Anthropic, Mistral and xAI alongside its own. The article borrows the venture capitalist Chris Dixon's framing of commoditising the complement, with Oracle's 2006 support for Linux as the historical analogue.

The security content is the substantive part. AWS describes three attack classes. The first is supply-chain poisoning with conditional triggers: adversaries are using AI not just to find vulnerabilities in open-source code but to poison libraries in ways other AI security programs do not detect — malware that executes only when a user issues a prompt containing a particular typo, or that activates only once other code is later added to the library. The second is trust-building through synthetic maintainers. Rick Anthony, Senior, who manages Amazon Inspector, describes attackers creating packages that deliver real, useful benefits, and notes that with AI coding agents they can maintain very reasonable-looking contribution histories and useful release cycles until they look like good citizens of the open-source community. The third targets the reviewers themselves: because more security review now runs through AI agents with their own blind spots, attackers try to fool the AI by giving it enough evidence to conclude that what it is looking at is fine.

On attribution, AWS chief security officer Stephen Schmidt names China, Russia and North Korea as sources of increasingly sophisticated attacks, and says China and Russia are well positioned because they face no penalties for running experiments against real-world targets. AWS runs red teams equipped with their own AI agents, both to find code vulnerabilities before adversaries do and to attack emerging open-weight models in order to discover adversary tactics first. Schmidt flags an asymmetry on the defensive side: finding a hole is only the first part, and developing a useful patch takes longer, with patches themselves tested for how they respond to adverse behaviour because adversaries will go after them the moment they ship. One capability data point anchors the trend — after Anthropic released its Mythos model to a handful of companies, the number of vulnerabilities researchers found and disclosed quickly doubled, and so did the number of patches. No specific CVEs, package names or malware families are disclosed; the poisoning techniques are described generically by AWS personnel rather than tied to published incidents.

supply chain open weights threat intelligence AWS
#4
Frontier LLMs 2026-08-03 Hacker News — AI front pageQwen Team 7.7 8.8/8.2/9.0 -1.0 frontier_llm

Alibaba's Qwen team released Qwen3.8-Max, a 2.4-trillion-parameter mixture-of-experts model with 95 billion active parameters built on the Qwen3.5 architectural foundation. It is available now through QwenCloud, and the team says open weights will land on Hugging Face and ModelScope next week — the first time a Max-class Qwen model has been open-weighted. The context window is one million tokens with a published effective-context figure of 95 percent, maximum output is 65,536 tokens, inputs are text and image, and reasoning effort is exposed at three levels. The API is protocol-compatible with both OpenAI chat completions and responses and with Anthropic's, and the post ships copy-paste configurations for Claude Code, Codex, Qoder, Qwen Code and OpenClaw.

On coding agents the picture is mixed rather than a clean sweep. Against Opus 4.8, Fable 5 and GPT-5.6 Sol at maximum effort, Qwen3.8-Max scores 86.6 on Terminal Bench 2.1 against 84.6, 84.6 and 88.8; 67.7 on SWE-bench Pro against 69.2, 80.0 and 64.6; 56.6 on DeepSWE 1.1 against 59.0, 70.0 and 73.0; and 73.5 on FrontierSWE against 70.0 and 88.8. It does lead on PaperBench at 93.0 against 80.3, 88.8 and 90.5. On general agentic and knowledge work it takes CoWorkBench at 74.8, WideSearch at 81.9, IFBench at 82.8 against a next-best 72.7, HealthBench at 60.2 and PRBench-Finance at 58.3, while trailing badly on Humanity's Last Exam at 43.6 against Fable 5's 53.3. Multimodal results include 82.3 on MMMU-Pro, 86.1 on OSWorld-Verified, 91.5 on Parametric CAD Bench, and 87.0 on Dense200 where Opus 4.8 scores 20.8.

The more interesting evidence is the long-horizon demonstrations. The team pointed the model at building a command-line tool from scratch and let it run autonomously; after roughly sixteen days the repository held 265 commits, 127 pull requests and 151 issues. Given only a paper on unified data selection for reasoning and access to GPUs, it ran about 125 hours over five days, wrote roughly 7,600 lines of code, took more than 1,100 actions and ran 33 rounds of GPU training — reproducing all six of the paper's findings in about 37 hours, then spending 88 more hours testing eighteen self-invented ideas across four rounds and lifting AIME24 from a 49.58 percent baseline to 52.29 percent. In an autonomous chip-design run it drove a crypto accelerator from 8,298 gates down to 678, with the single largest step coming from replacing a modulo divider with an iterative shift-subtract; after place and route the die shrank from 106 by 106 micrometres to 46 by 46, wirelength fell from 33,369 to 4,187 micrometres, and timing closed at 500 megahertz with positive slack where the initial design had been 4.46 nanoseconds short.

The benchmark caveats are substantial and mostly stated in the post's own footnotes. A large share of headline numbers come from in-house benchmarks where Qwen sets the harness, the judge and the scoring. Harnesses are not held constant across models — Qwen is evaluated on Claude Code, OpenCode or Qwen-Agent depending on the benchmark while competitors are scored on Terminus 2, Codex, or the best published score across harnesses. SWE-bench Pro results use a self-refined variant with problematic tasks corrected, several ground truths were manually corrected by Qwen, and PaperBench is judged by Claude Opus 4.6. Fable 5 results may involve fallbacks. No pricing is disclosed anywhere in the post, and the open weights are promised but not yet shipped.

Qwen MoE agentic coding open weights
#5
Industry 2026-04-24 Hacker News — AI front pageModel Republic 7.6 7.0/7.6/8.2

An investigation published by Model Republic and resurfaced on the Hacker News front page overnight documents Acutus, an anonymously operated news site that launched on 29 December 2025 and published 94 full-length articles in under four months with no masthead, bylines, named editors or ownership disclosure. Run through Pangram, an AI content detector whose vendor claims a near-zero false-positive rate, 69 percent of those articles flagged as fully AI-generated and 28 percent as partially AI-generated; only three classified as human-authored. The trigger was an interview request sent to Encode's general counsel by a reporter named Michael Chen writing from a generic address and offering only a written question-and-answer format; that email also came back fully AI-generated.

The technical reconstruction is the strongest part. The site is a React application whose client-side JavaScript bundle exposes the internal editorial interface, including an article-creation form with fields labelled AI Background Context and Question Prompts, a Generate Story Draft button and a Regenerate button, plus tools to extract quotes from research notes and run a multi-pass scored editorial review. A public endpoint that populates the homepage returns the full story database plus the production record: five review categories, four scored out of a hundred covering style compliance, quote accuracy, source verification and one internal metric, along with fact-checking status, flagged issues, proposed corrections and timestamped resolutions. Those timestamps show the entire multi-pass review completing in a median of 44 seconds from first issue resolved to last, with the publish click landing a median of ten seconds after the final resolution. Forty-two of the 94 stories carry an automated reviewer status of needs_revision and were published anyway. One flagged field preserves the model's original wording alongside the suggested edit.

The infrastructure is built for machine consumption: permissive AI-crawler rules, a deprecated ChatGPT plugin manifest, and an experimental llms.txt describing the site as independent journalism operating to a strict zero-hallucination editorial standard. All content is Creative Commons licensed and syndicated as a wire service over RSS.

The funding chain is where the reporting turns circumstantial, and the author says so. Acutus has almost no public footprint — four links on Twitter in total, two of them written or retweeted by the president of a Republican public affairs firm. Coverage overlaps that firm's client list in several places, including a piece attacking pharmacy benefit managers ten days before related reform was signed into law, whose internal source log lists statements from a trade group that is a client but who is never quoted. The third name on that firm's client list is a political consultancy whose chief executive co-founded a $125 million super PAC funded primarily by OpenAI's president alongside an OpenAI investor and assembled under the company's chief political operative. The headline says appears to be, the body says may be, and no direct financial documentation is presented — the chain rests on a client list, retweet patterns and topical overlap. The findings also depend on a single commercial detector, and some quotes come from real people who were genuinely approached, so the authorship claim is narrower than a claim that all quotes are fabricated. The article does not state whether any of the named parties were contacted for comment.

AI-generated content detection media
#6
Infrastructure 2026-07-31 Hacker News — AI front pageWafer 7.5 7.5/7.2/7.8

Wafer published a serving study putting Kimi K3 — 2.8 trillion parameters, over 1.5 terabytes of VRAM for weights alone before any key-value cache for a million tokens of context — on AMD MI355X and comparing against NVIDIA B200 and B300 nodes. The memory arithmetic dictates the deployment shape: a B200 node with 192 gigabytes per GPU cannot fit K3 plus a one-million-token cache pool, forcing either a B300 node at 288 gigabytes per GPU or two B200 nodes at tensor-parallel sixteen. MI355X also carries 288 gigabytes and is roughly 2.4 times cheaper per GPU than a B300 and 1.7 times cheaper than a B200.

At 1,024 tokens in and 400 out, the three configurations give per-stream decode of 118, 90 and 172 tokens per second for MI355X, the two-node B200, and B300 respectively; peak aggregate of 952, 498 and 1,568; and peak aggregate per GPU of 119, 31 and 196. Normalised by assumed rental prices of $2.50, $4.25 and $6.00 per GPU-hour, the per-dollar figures are 48, 7 and 33 tokens per second per dollar. So MI355X delivers more than 3.8 times the aggregate node throughput and more than 1.3 times the single-stream decode of the two-node B200 setup, while the B300 wins outright on absolute throughput by about 1.65 times at 2.4 times the price. The B200 numbers are somewhat deflated by being the only two-node configuration, paying a cross-node all-reduce on the decode critical path over RoCE version two at roughly 195 gigabits per second.

Two engineering findings carry beyond the price comparison. Kimi K3 ships zero draft tensors — no multi-token prediction, no EAGLE head — so the only speculative path is an external block-diffusion draft model, which worked on CUDA but crashed on ROCm with an undefined top-k renormalisation symbol. The cause was that SGLang's accept-sampling verifier has a dense path calling a kernel imported from the CUDA build, while the ROCm build aliases only a Triton top-p kernel, leaving the top-k function undefined on the target architecture and taking the scheduler down. The fix was a plain PyTorch sort, masked fill and divide dropped into the ROCm sampling branch — no custom kernel required — and speculative decoding then delivered roughly 2.2 times single-stream, about 1.7 times per-stream at moderate load, and eighteen percent on peak aggregate, with the aggregate peak shifting to much higher concurrency.

The second was prefill. An identical 172,000-token cold prefill took roughly 51 seconds on MI355X against 23 on a B300, because the fast AITER multi-head latent attention prefill kernel was failing to load: K3 at tensor-parallel eight yields twelve attention heads per rank, and AITER's path supports four, eight, or multiples of sixteen. Zero-padding the head count from twelve to sixteen, running the fast kernel and extracting the real twelve heads took the AITER assembly path to roughly 13,000 tokens per second steady-state against the Triton fallback's four to seven thousand — a two-to-threefold prefill speedup. The authors are careful that this is a time-to-first-token lever, not an aggregate-throughput one, and it does not move any headline number. The benchmark is a single sequence shape with no accuracy checks, and the claim that AMD's persistent software gap is closing is presented as Wafer's thesis rather than a measured result.

AMD MI355X serving SGLang speculative decoding
#7
Research 2026-07-26 Hacker News — AI front pagearXiv cs.CL (Computation & Language) 7.5 7.4/7.8/7.2

A team spanning Stony Brook computer science, Columbia Law, Michigan and the MIT Initiative on the Digital Economy assembled 14,419 randomly selected self-published genre-fiction e-books released between January 2023 and March 2026, with daily sales observed through June 2026, and paired them with a proprietary daily panel from a major publisher covering roughly 500,000 Amazon identifiers and an estimated 95 percent or more of daily unit volume. Full texts were obtained lawfully through author emails, library loans and purchases. None of the books disclose AI content. Every chapter was scored with Pangram version 3.3, and titles were binned into no AI text, light AI at 25 percent or less of analysed windows, and substantial AI above 25 percent.

Books with substantial AI text make up 20.0 percent of titles but only 12.1 percent of sales and 11.3 percent of revenue, while no-AI books are 62.9 percent of titles and 71.7 percent of sales. The light band sits near-neutral at 17.1, 16.2 and 16.2 percent, so the shortfall is specific to the substantial band. In the upper tail the gap widens: among the top five percent by launch-window sales, no-AI share rises from 63 to 73 percent and substantial-AI falls from 20 to 10 percent.

The dilution arithmetic is the paper's central number. Indexed to the first quarter of 2023, the cumulative released catalogue grew 38.3 times and the number of books selling in a quarter grew 19.2 times, while quarterly unit sales grew only 7.3 times and revenue 8.9 times. Launch-window revenue per selling title fell from the 2023 to the 2025 release cohort in six of eight genre clusters for all books — and in seven of eight when restricted to no-AI books alone, which rules out a pure mix-shift explanation. Only fantasy, supernatural and horror rose. Across genre-months, the no-AI share of top-25 rank slots falls from about 88 percent in the lowest substantial-AI-exposure bin to about 63 percent in the highest, and in genres where Kindle Unlimited is heavily used the no-AI lead over substantial-AI is 8.5 percentage points smaller in sales share.

Two further results are worth flagging. Production is concentrated: of 824 authors in the event-time panel, 311 raised monthly output after adoption, the top quartile produced 60.5 percent of post-adoption substantial-AI books, and the top individual substantial-AI title grossed $643,000 on 80,431 sales. And a rare-expression overlap measure — five-word-or-longer expressions appearing in five or fewer Google Books volumes and absent from a 4.7-trillion-token web snapshot — shows substantial-AI bestsellers at 45.0 percent token-level coverage against 37.7 for no-AI in the top fifty by revenue, with coverage rising 7.6 percentage points per tenfold revenue increase for AI books against 1.1 for human ones. A reference set of 200 literary award winners sits at just 19.1 percent.

The authors are careful about what this supports. Every comparison is observational, built on release cohorts and genre-level exposure rather than a causal design, and the Kindle Unlimited result is explicitly labelled heterogeneity rather than a causal effect. All labels depend on one commercial detector with no independent validation on this corpus. Author identity is byline-level, so pen names understate concentration. The sales panel does not separate page-reads from purchases, which is a real gap given the subscription finding, and revenue is gross consumer spending before platform commissions. The overlap measure shows aggregate similarity to prior books and explicitly cannot attribute any passage to a specific title. The framing is aimed squarely at the market-effect prong of fair use and the market dilution theory raised in recent litigation.

cs.CL market dilution copyright detection
#8
Robotic Autonomy 2026-08-03 Hugging Face Daily PapersAK (@_akhaliq) Daily Papers 7.4 6.6/6.4/6.3 +1.0 robotic_autonomy

N_0-VTLA adds a tactile pathway to a vision-language-action backbone through a three-stage recipe: visuo-tactile pre-training on NeoData, a large-scale visuo-tactile robot corpus; staged integration of a predictive tactile pathway that distills contact priors into fine motion adjustments; and ALTER, an advantage-conditioned offline RL method that converts relative progress and trajectory-event comparisons into binary advantage labels for training on a fixed deployment corpus. The authors claim it is the first VLA pretrained on tactile data at scale, targeting contact-rich manipulation including deformable objects.

cs.RO VLA tactile
#9
Government & Defense 2026-08-03 AK (@_akhaliq) Daily PapersarXiv cs.CV (Computer Vision)arXiv — Evals & BenchmarksarXiv — AI, Defense & National Security 7.3 6.2/6.5/6.2 +1.0 gov_defense

An audit of SULAND, the main public RGB benchmark for PFM-1 and PMA-2 surface-landmine detection from UAV and UGV imagery, found missing and false annotations, localisation errors, inconsistent visibility criteria and an inverted out-of-distribution class-ID convention. SULAND_v2 preserves the original images and splits but manually revises annotations, yielding 33,771 images and 12,433 boxes, and benchmarks 35 detector configurations across nine families. Annotation refinement alone lifts YOLOv8 in-distribution test mAP@50 by 14.6 to 19.6 percentage points, and fixing the class-ID convention raises mean OOD mAP — a reminder that in safety-critical detection, label quality dominates architecture choice.

cs.CV object detection OOD
#10
Robotic Autonomy 2026-08-03 arXiv cs.RO (Robotics)arXiv — Generative Media / DiffusionarXiv — Evals & BenchmarksarXiv — Robotic Autonomy / Embodied AI 7.2 6.2/6.3/6.0 +1.0 robotic_autonomy

Rather than bolting a safety filter onto a policy's final output, this modular inference framework modifies the flow-matching denoising process itself so generated trajectories are safe by construction. A smooth log-sum-exponential aggregate barrier enforces the constraint over entire action chunks with minimal compute overhead and without altering semantic intent, and the authors bound the 2-Wasserstein distance between generated and target distributions. No safety-specific dataset or retraining is required; validation covers two manipulation platforms and a 2D navigation benchmark with no measured success-rate degradation.

cs.RO CBF flow matching safety
#11
Robotic Autonomy 2026-08-03 arXiv cs.RO (Robotics)arXiv — Reinforcement LearningarXiv — Evals & BenchmarksarXiv — Robotic Autonomy / Embodied AI 7.2 6.3/6.2/6.0 +1.0 robotic_autonomy

Critic-based RL post-training for VLA policies usually estimates value from single-frame observations or single-frame VLM latents, which mismatches the partial observability of robot control; naively feeding observation history in is exponentially expensive and still fails because scalar-return regression gives too little supervision to learn temporal dynamics. WCM diagnoses this as a state-approximation problem and builds a lightweight LeJEPA critic that jointly predicts future latent state and estimates value, so the representation is explicitly trained on temporal structure. It drops into both on-policy and off-policy pipelines.

cs.RO VLA RL world model
#12
Robotic Autonomy 2026-08-03 arXiv cs.RO (Robotics)arXiv — Robotic Autonomy / Embodied AI 7.1 6.4/6.2/5.8 +1.0 robotic_autonomy

Managed supervised fine-tuning APIs are becoming the access route to closed-weight robot foundation models: you submit data and receive a tuned policy, with no weights, gradients or training internals. That restricts improvement to pure imitation and rules out closed-loop methods. CLIFT works inside those constraints, running non-invasive closed-loop iterative fine-tuning against Gemini Robotics On-Device to produce humanoid specialists — effectively recovering some of the benefit of RL through repeated data curation rather than through gradient access.

cs.RO humanoid fine-tuning API
#13
Robotic Autonomy 2026-08-03 arXiv cs.RO (Robotics)arXiv — Generative Media / DiffusionarXiv — Robotic Autonomy / Embodied AI 7.1 6.2/6.0/6.0 +1.0 robotic_autonomy

FibVLA targets the tension between capturing temporal history and keeping VLA inference fast. It applies logarithmic hindsight sampling to both proprioceptive states and visual frames so long-term dependencies are covered with minimal redundancy, uses flow matching in the action expert to produce action distributions, and adds a Fibonacci recurrent inference schedule that generates long-range planning steps under real-time closed-loop feedback. Reported gains are in action smoothness and success rate with no retraining of the large visual encoders.

cs.RO VLA efficiency
#14
Robotic Autonomy 2026-08-03 arXiv cs.RO (Robotics)arXiv — Evals & BenchmarksarXiv — Robotic Autonomy / Embodied AI 7.0 6.2/6.0/5.9 +1.0 robotic_autonomy

Physics simulators demand asset construction and calibration and still leave a sim-to-real gap, while video generators lack precise control over responses to fine-grained robot actions. The Boundless World Model combines initial-environment guidance, dynamic visual history and temporally aligned action conditioning for stateful autoregressive prediction of future observations, explicitly including risky and failure-prone outcomes that are expensive to collect on hardware. It is released open-source.

cs.RO world model simulation
#15
Government & Defense 2026-07-22 Cognition AI (Devin) 7.0 5.8/6.5/5.7 +1.0 gov_defense

Cognition has signed a memorandum of understanding with the US Department of Energy to join the Genesis Mission, a national initiative the company describes as America's Manhattan Project for AI. Separately in the same window it announced a partnership deploying Devin across a cybersecurity practice serving more than 260 clients including 26 of the Fortune 500 and the top five global banks. Both were captured from the blog listing this run; neither had appeared in a prior digest window.

DOE coding agents partnership
#16
Generative Media 2026-08-03 AK (@_akhaliq) Daily PapersHugging Face Daily PapersarXiv cs.CV (Computer Vision)arXiv — Efficiency (Quantization, MoE, Inference)arXiv — Evals & Benchmarks 7.0 7.2/6.9/6.9

Text conditioning has resisted scaling analysis because diffusion loss does not scale with prompt token count. The finding here is that converged diffusion loss does scale with the amount of structured language in the prompt, quantified two ways: a white-box likelihood metric and a black-box attribute metric. Across controlled runs the converged loss decreases roughly linearly in the first and follows a power law in the second. Acting on that, the authors improve diffusability by building structured prompts with semantic and geometric annotations derived from images, and promptability by training a prompter through supervised fine-tuning, cold start and verifier-gated on-policy distillation. The system beats all evaluated open-weight models on nearly every compositional, reasoning and world-knowledge benchmark.

How it was discussed
  • AK's Daily Papers thread emphasised the scaling-law framing; Hugging Face Daily Papers foregrounded the resulting open-weight benchmark sweep.
  • The efficiency feed picked it up for the prompter distillation recipe rather than the conditioning result.
cs.CV diffusion scaling laws
#17
Reinforcement Learning 2026-08-03 AK (@_akhaliq) Daily PapersHugging Face Daily Papers 7.0 7.1/7.0/6.8

RLVR's reach stops where correctness stops being deterministically checkable, and the usual substitutes — human preferences, reward models, LLM judges — bring evaluation bias, judge capability ceilings and inference cost. Borrowing the pretext-task idea from self-supervised learning, RLSVR transforms open-ended tasks into verifiable proxy environments whose internal rules generate reward automatically. The instantiation, SpyRL, gives agents asymmetric information on a shared target task and has them vote to identify a designated spy; because the spy identity is predetermined, voting outcomes are fully verifiable rewards.

How it was discussed
  • Both Hugging Face Daily Papers and AK's thread framed the contribution as extending RLVR's reach rather than as a self-play result.
cs.CL RLVR self-play
#18
Safety, Policy & Regulation 2026-08-03 AK (@_akhaliq) Daily PapersHugging Face Daily Papers 7.0 7.1/7.4/6.6

Safeguards decide whether to answer before seeing how the answer will be used, which is a structural problem for dual-use tasks: the same answer helps an authorised professional or an attacker, and an attacker can imitate a benign request and interaction history. Separating the capability the model releases from the evidence available about downstream use, the authors derive an exact worst-case floor on attacker assistance conditional on that evidence being copyable, yielding a trilemma between useful capability, reliable safety and open access. They then show a trusted credential — hard-to-copy information predicting actual downstream use — can complement existing safeguards, and identify the stronger condition needed to eliminate the floor entirely.

cs.CL dual use safeguards
#19
Robotics 2026-08-03 arXiv cs.RO (Robotics)arXiv — AI for Science 7.0 5.8/6.5/5.8 +1.0 robotics

A capability assessment of humanoid and quadrupedal platforms across five axes — hardware, locomotion, autonomy, data and applications — identifying recent advances and the open challenges blocking widespread adoption. The outlook section covers ethical considerations, economic potential, policy implications and broader societal effects, and treats scientific discovery as a distinct use case rather than a byproduct of commercial deployment.

cs.RO survey humanoid quadruped
#20
Robotic Autonomy 2026-08-03 arXiv cs.RO (Robotics)arXiv — Generative Media / DiffusionarXiv — Robotic Autonomy / Embodied AI 7.0 6.1/5.9/5.9 +1.0 robotic_autonomy

Standard diffusion and flow-matching policies couple an uninformative Gaussian prior to the action space, forcing the model to learn a complex, high-cost vector field. Temporal Policy reformulates action generation as a temporally coupled transport problem using stochastic interpolants, initialising the flow at the robot's recent history so past states are explicitly coupled to future action sequences. The data-dependent coupling reduces transport cost and produces straighter fields, which is where the inference-cost saving comes from.

cs.RO flow matching LfD
#21
Agents & Tool Use 2026-08-03 AK (@_akhaliq) Daily PapersHugging Face Daily PapersarXiv cs.AI (Artificial Intelligence)arXiv — Agents / Tool UsearXiv — Evals & Benchmarks 6.9 7.0/6.8/6.8

ExtractBench covers 4,869 pages across 370 enterprise documents, 8 business domains and 67 document types, with ground truth built from independent-system agreement on real documents, known values for synthetic lists and human verification for forms. It is the first benchmark the authors know of to score value accuracy, record completeness at scale, grounding and measured cost together, reporting order-insensitive value F1 plus word- and page-level grounding F1. The headline pattern: commercial VLMs do well on short documents but truncate record lists on long ones, while coding agents hold accuracy at much higher cost.

How it was discussed
  • Hugging Face Daily Papers surfaced it primarily as an enterprise-agent evaluation rather than a document-AI result.
  • The arXiv agents and evals feeds both picked it up, reflecting that the contribution is as much a cost metric as an accuracy one.
cs.AI benchmark document AI
#22
Robotic Autonomy 2026-08-03 arXiv cs.RO (Robotics)arXiv — Agents / Tool Use 6.9 6.0/5.9/5.8 +1.0 robotic_autonomy

Training-free vision-and-language navigation queries a multimodal LLM each step, and long horizons turn either image streams or dense maps into a memory and reasoning bottleneck. HAM-VLN makes memory decision-coupled and agent-authored: in the same model call that selects the next action, the robot records room type, objects, navigation progress and failure notes into a persistent depth-grounded world graph, keeping recent waypoints verbatim inside a bounded window while older context is compressed into the graph.

cs.RO VLN memory
#23
Robotic Autonomy 2026-08-03 arXiv cs.RO (Robotics)arXiv — Efficiency (Quantization, MoE, Inference)arXiv — Evals & Benchmarks 6.9 5.9/5.8/5.9 +1.0 robotic_autonomy

End-to-end driving models are black boxes that struggle in complex scenes, and the usual fix — bolting on a VLM for explicit reasoning — depends on pre-generated annotations that are both expensive and potentially wrong. This teacher-student framework has the teacher VLM generate logical explanations and then reflectively refine them against outcomes, so supervision is grounded in driving results rather than in human-labelled rationales, combining structured reasoning with geometric precision in the student.

cs.RO autonomous driving distillation
#24
Post-Training 2026-08-03 arXiv cs.LG (Machine Learning)arXiv — Efficiency (Quantization, MoE, Inference)arXiv — Reinforcement LearningarXiv — Evals & Benchmarks 6.9 7.0/6.9/6.7

RLVR broadcasts one response-level reward across every token; on-policy distillation scores each token against a stronger teacher for a dense advantage but caps the student at teacher quality. Combining them with a fixed coefficient triggers entropy collapse through two miscalibrations — a magnitude mismatch where token-level OPD advantages spike past the bounded RLVR advantage and erase it, and a temporal mismatch where sustained full-strength OPD keeps pulling the student toward the teacher. SAF applies sparsify-then-compress for magnitude and warm-up-then-anneal for time, only to the OPD advantage, evaluated with GRPO across seven math and code benchmarks on Qwen3 at 1.7B, 4B and 8B.

cs.LG RLVR distillation
#25
Safety, Policy & Regulation 2026-08-03 AK (@_akhaliq) Daily PapersHugging Face Daily Papers 6.8 6.7/7.0/6.6

System prompts govern foundation-model behaviour throughout commercial products but are rarely disclosed to the public or to regulators. AISPA evaluates individual instructions along eight user-relevant dimensions, classifying each as protective or problematic, applied to 3,249 instructions drawn from 88 commercial products. Design varies enormously — some organisations average over 60 protective instructions per product while others average fewer than five — and protective coverage is wide but shallow: 98.9 percent of products contain at least one protective instruction, yet only 24 percent cover all eight dimensions. System prompts have also grown steadily longer and more protective over time.

cs.CL auditing transparency
#26
Robotic Autonomy 2026-08-03 arXiv cs.RO (Robotics) 6.8 5.9/5.8/5.6 +1.0 robotic_autonomy

Driving world models typically do dense prediction of future video, occupancy, bird's-eye-view representations or agent motion. Auto-JEPA argues planning only needs the scene features that affect future ego action, and predicts an intent embedding aligned with the latent representation of the future ego trajectory from visual observations, egomotion history and navigation commands. The predicted intent retrieves executable trajectories from a fixed memory, which are then ranked by a scene-conditioned scorer.

cs.RO JEPA world model driving
#27
Generative Media 2026-08-03 AK (@_akhaliq) Daily PapersHugging Face Daily PapersarXiv cs.CV (Computer Vision)arXiv — Generative Media / DiffusionarXiv — Reinforcement Learning 6.8 6.9/6.8/6.6

RL works well for text-to-image and single-image editing but does not transfer to multi-reference editing because no reward model captures multi-image relational constraints, and using an MLLM zero-shot forces a choice between hallucination-prone long-form reasoning and under-powered short-form judgments. EVR decomposes evaluation into distinct visual criteria; per criterion an Evaluator generates multiple candidate hypotheses and a Verifier grounds each claim in concrete visual evidence before accepting it. That yields fine-grained rewards usable for RL fine-tuning of off-the-shelf editors with no architectural change, with substantial gains over base Qwen-Image-Edit.

cs.CV RL image editing
#28
AI Coding 2026-07-30 Hacker News — AI front pagearXiv cs.SE 6.8 6.7/6.7/7.0

A teacher-student loop pairs a deterministic migrator, analyser, runner and parity gate with an agentic authoring layer that may only propose mutations, never accept them. Witness Search runs six input-space algorithms in parallel — pairwise and three-way interaction testing, Latin hypercube sampling, adaptive random testing, MAP-Elites and UCB1 — all converging within two to three branches of each other, which the authors read as a structural ceiling. When search plateaus, the authoring layer force-sets a mock-returned value symmetrically in both COBOL and Java. On a 4,114-line production-shaped program, coverage reached 135 of 142 paragraphs and 91.9 percent of branches across 166 executions with zero parity failures. The load-bearing caveat is the authors' own: the oracle is the legacy program, so the loop certifies compatibility rather than correctness, and migrated bugs are reproduced by design. No model is named.

cs.SE migration agents
#29
AI for Science 2026-08-03 arXiv cs.AI (Artificial Intelligence)arXiv — Agents / Tool UsearXiv — AI for SciencearXiv — Evals & Benchmarks 6.8 6.9/6.8/6.6

Long-horizon scientific analysis has lacked process-supervised environments over real scientific data, so agents get rewarded on final claims rather than on the analytical steps that produce them. SciDisco compiles hypotheses, datasets, hidden evidence graphs and verifiers into task environments where progress is checkable mid-interaction, synthesises verifier-filtered multi-turn demonstrations grounded in a DAG, and uses DiscoPO to assign turn-level credit to actions producing verifiable evidence. SciDisco-14B is reported at state of the art on hypothesis-driven data-analysis benchmarks.

cs.AI agentic RL scientific discovery
#30
Industry 2026-07-30 Hacker News — AI front pageBusiness Insider 6.7 6.3/7.1/6.7

Apollo Global Management analysts scored 321 US occupations against Anthropic's Economic Index, which measures the share of an occupation's tasks observed being performed with Anthropic tools, and compared Bureau of Labor Statistics wage data for 2022 against 2024. The most AI-exposed occupations show an average 6.7 percent decline in real wage growth after 2023, with no detectable effect on employment levels. The distribution is skewed: service workers average a 24.3 percent decline in earnings growth, the bottom quartile of earners 10.7 percent, and there is no significant effect among the highest paid. The exposure measure is one vendor's product telemetry, the window is short, and the data contains visible confounds — radio DJs show a 52 percent real wage collapse at low exposure while administrative law judges gained 17.5 percent at moderate exposure.

labour economics measurement
#31
Post-Training 2026-08-03 AK (@_akhaliq) Daily PapersHugging Face Daily Papers 6.7 6.8/6.7/6.5

Rubric-based RL is known to suffer from Unexplored Criteria that no rollout satisfies and therefore receive no gradient; the usual fix injects rubric text as rollout guidance, creating a train-inference mismatch. This paper names a second failure mode: Suppressed Criteria, satisfied by some rollouts but whose signal is lost because scalar reward aggregation assigns them non-positive aggregate advantage. Over 57 percent of samples show this throughout training, averaging 1.8 suppressed criteria each. Criterion-Distilled Policy Optimization addresses both without the guidance mismatch.

cs.CL rubric RL GRPO
#32
Robotic Autonomy 2026-08-03 arXiv cs.RO (Robotics) 6.7 5.8/5.7/5.5 +1.0 robotic_autonomy

Rule-based multi-robot coordination depends on predefined task models and bespoke decision programs, which do not survive complex semantic instructions or heterogeneous platforms. D-VLC pushes instruction interpretation, task decomposition and role assignment into a decentralised vision-language layer so each robot reasons about objects, regions and spatial relationships locally, without a central planner holding a global model of the environment.

cs.RO multi-robot VLM
#33
Robotic Autonomy 2026-08-03 arXiv cs.RO (Robotics)arXiv — Evals & Benchmarks 6.7 5.8/5.7/5.6 +1.0 robotic_autonomy

RGB observations carry no explicit geometry, so visuomotor policies learned from them break under camera perturbation. RayViT represents camera geometry as a Plucker ray map, patchifies it into ray features, and uses gated cross-attention to build a ray-conditioned class token; ray features are added as dense positional embeddings and the ray class token replaces the original ViT class token. The base weights stay pretrained, so the cost is a small conditioning module rather than a new encoder.

cs.RO ViT imitation learning
#34
Post-Training 2026-08-03 arXiv cs.CL (Computation & Language)arXiv — Reinforcement LearningarXiv — Evals & Benchmarks 6.7 6.8/6.6/6.6

Translation with Thought trains a resource-rational policy that modulates inference between intuitive and deliberate reasoning according to domain difficulty: supervised fine-tuning on difficulty-aware long chain-of-thought traces distilled from DeepSeek-R1 and rewritten for reasoning economy, then RL with a hybrid reward balancing translation quality against reasoning cost. Across 15 benchmarks spanning in-domain and out-of-domain settings, 3 seen and 59 unseen languages, TwT at 7B and 14B outperform much larger state-of-the-art reasoning models while cutting token usage by 32 to 60 percent.

cs.CL translation reasoning efficiency
#35
Robotics 2026-08-02 Two Minute Papers 6.7 5.6/5.8/5.7 +1.0 robotics

A walkthrough of NVIDIA-affiliated work on human-to-humanoid transfer arguing that pure motion imitation fails because the human demonstration does not encode the contact forces and balance corrections the robot actually needs. The video links to the project page for the underlying method; the episode is a summary rather than a primary release.

humanoid imitation learning video
#36
Evaluations & Benchmarks 2026-08-03 arXiv cs.AI (Artificial Intelligence)arXiv — Agents / Tool UsearXiv — Evals & Benchmarks 6.6 6.7/6.7/6.3

Analysing 22,500 agent trajectories across GAIA, SWE-bench and Multi-Challenge, the authors define an Evaluative Dissonance Index quantifying how much LLM-as-a-judge systems mistake structural proceduralism for semantic truth when a candidate mimics consensus. A logistic meta-evaluator isolates the syntactic triggers of evaluator capture at 0.8779 ROC-AUC, and zero-shot leave-one-domain-out transfer at 0.7482 mean ROC-AUC indicates the vulnerability is domain-agnostic rather than an artifact of any one benchmark.

cs.AI LLM-as-judge evaluation
#37
Industry 2026-08-01 Hacker News — AI front page 6.6 6.4/6.2/7.2

Two Google product retractions landed within a day of each other and both reached the Hacker News front page: the AI Studio app was cancelled after roughly 800,000 preorders, and a separate Earth AI generator was withdrawn about a day after launch. Announcements came through the respective product accounts rather than a formal blog post, so the stated reasoning is thin; the pairing is what drove the community reaction.

Google product launch
#38
Evaluations & Benchmarks 2026-07-21 Hacker News — AI front pageMIT Sloan 6.6 6.4/6.8/6.5

Researchers built a life-cycle model as a normative benchmark, had 1,000 adults write their own advice-seeking prompts, and simulated ages 22 to 89 following GPT-5.2, GPT-5.6 and Gemini 3 Flash recommendations. Advice was broadly sound — saving during working years, drawdown in retirement, diversified equity, de-risking after 45 — but handled shocks poorly, leaned on rules of thumb and let portfolios drift instead of rebalancing. Prompts written by men, by more financially literate users, or by users with prior AI experience produced roughly 5 percent more wealth near retirement, compounding to about $50,000 less at age 60 for women and less literate users; about two-thirds of the gender gap traces to how prompts are written and one-third to the model changing advice when the same prompt is labelled as coming from a woman. Models also volunteered specific providers users never mentioned — Vanguard in 6 percent of responses, iShares in 3.4, against under 0.4 percent of prompts naming either.

evaluation bias finance
#39
Generative Media 2026-08-03 AK (@_akhaliq) Daily PapersarXiv cs.CV (Computer Vision) 6.6 6.7/6.5/6.5

Mainstream mesh generation serialises a mesh into tokens and decodes autoregressively, which is slow and accumulates error. Meshy T2 uses a vertex-set VAE that encodes a mesh into one continuous latent token per vertex and decodes vertices, edge connectivity and face winding order in a single pass — no vertex quantisation, no welding, artist-authored topology preserved. Generation is a coarse-to-fine cascade: an image-conditioned voxel flow sketches an occupancy scaffold, then a mesh flow populates it with per-vertex latents conditioned on image, scaffold and a requested vertex budget, giving explicit face-count control.

cs.CV 3D flow matching
#40
Safety, Policy & Regulation 2026-08-02 The Cognitive Revolution (Nathan Labenz) 6.6 6.3/6.9/6.5

The second instalment of the China travel series covers how AI safety work is organised and practised inside Chinese labs and institutions, following a first part on the broader research and deployment landscape. Useful primarily as first-hand reporting on institutional structure and evaluation practice, an area where secondary coverage tends to rely on translated policy documents rather than direct conversation.

podcast China governance
#41
Reinforcement Learning 2026-08-03 arXiv cs.CL (Computation & Language)arXiv — Reinforcement LearningarXiv — Evals & Benchmarks 6.6 6.7/6.6/6.4

Long chain-of-thought reasoning brings redundancy accumulation, context overflow and error anchoring, and the authors argue the real bottleneck under a bounded window is neither trajectory compression nor test-time control but the absence of a reusable intermediate interface that can replace discarded history. They also identify a specific failure of outcome-reward long-chain RL: when the model has not solved the task and the window is nearly exhausted, the final-answer reward rewards premature guessing. ThinkReset instantiates interface writeback and reset in text space and optimises post-reset continuation success directly.

cs.CL long horizon context
#42
Agents & Tool Use 2026-08-03 arXiv cs.AI (Artificial Intelligence)arXiv — Agents / Tool UsearXiv — Evals & Benchmarks 6.6 6.8/6.6/6.4

Most agent memory systems spend extra LLM calls generating intermediate records and mediating retrieval, adding recurring token and latency cost while merged or omitted details obscure the original evidence. Zero-Mem asks whether structured memory access needs generation at all: it preserves raw interaction traces and organises them two ways — an entity-context graph for cross-interaction connections and a temporal hierarchy for conversational locality — weighing both views per query. Deterministic calibration discards conflicting evidence, and only the final reader invokes a model.

cs.AI memory retrieval
#43
AI for Science 2026-08-03 arXiv cs.AI (Artificial Intelligence)arXiv — AI for SciencearXiv — Evals & Benchmarks 6.5 6.6/6.5/6.4

Evaluating autonomous research systems is unresolved because the output is a paper, not a score. This benchmarking protocol runs an automated peer-review system built on frontier models, assessing originality, scientific rigour, clarity and significance, and applies it to Sakana AI v1 and v2, CycleResearcher and Data-to-Paper. Each framework ran on a consistent set of 15 research proposals published by a commercial autonomous-science company, producing 60 papers for comparison under identical review conditions.

cs.AI autonomous research evaluation
#44
AI for Science 2026-08-03 arXiv cs.LG (Machine Learning)arXiv — AI for Science 6.5 6.6/6.6/6.3

I-MLFFs replace explicit stacks of neural network layers with self-consistent fixed-point equations, which lets intermediate representations be reused across successive molecular-dynamics timesteps to warm-start force evaluation. The result combines the compute footprint of a shallow single-layer force field with the representational capacity of a deep network, and the efficiency gain is architecture-agnostic — demonstrated across invariant, equivariant Cartesian and other graph neural network classes — because it comes from coupling force prediction to trajectory integration rather than treating them separately.

cs.LG molecular dynamics force fields
#45
Safety, Policy & Regulation 2026-08-01 TechCrunch — AI 6.5 6.2/6.9/6.4

A federal judge declined xAI's request to enjoin Minnesota's statute banning applications that let users generate nudified images of real people, allowing the ban to take effect while the company's lawsuit proceeds. The ruling is notable because it leaves a state-level restriction on a specific generative-image capability operating against a frontier lab, rather than against a downstream app developer.

litigation generative media state law
#46
Efficiency 2026-08-03 arXiv cs.LG (Machine Learning)arXiv — Post-training / AlignmentarXiv — Reinforcement Learning 6.5 6.6/6.5/6.3

Where LoRA adds a low-rank update to weight matrices, LARA reads the hidden state at a small set of layers and adds a low-rank correction back into the residual stream, leaving base weights untouched entirely. It matches LoRA at equal parameter count on a code fine-tuning task and on DPO. The practical difference is that because adaptation is a frozen base plus a residual, LARA exposes an inference-time scale that interpolates smoothly between base and adapted behaviour — graded control that weight-space adaptation does not offer, and a route to composing several adapters.

cs.LG PEFT LoRA
#47
Post-Training 2026-08-03 arXiv cs.LG (Machine Learning)arXiv — Post-training / AlignmentarXiv — Reinforcement Learning 6.5 6.6/6.5/6.3

Scalar reward models are efficient and probabilistically interpretable but lean on superficial cues that fail out of distribution; generative reward models reason their way to robustness but produce natural-language scores lacking numerical flexibility. Recent hybrids combine both through off-policy multi-task learning, which does not guarantee the generated reasoning actually helps the scalar head. This work trains the latent reasoning trace end-to-end against the scalar prediction so the trace is optimised for the downstream reward rather than alongside it.

cs.LG reward models RLHF
#48
Efficiency 2026-08-03 arXiv cs.LG (Machine Learning)arXiv — Efficiency (Quantization, MoE, Inference) 6.5 6.6/6.5/6.3

Multi-model systems increasingly share contexts, retrieved evidence and dialogue histories, but key-value caches are model-specific, so every model re-prefills or separately stores the same context. MoT translates a source model's context cache into a target model's cache space using multiple translator modules to capture diverse source-target mappings, rather than relying on a single projection path or a global shared latent space — the scaling argument being that one projection cannot cover architecturally distant pairs.

cs.LG KV cache multi-agent
#49
Agents & Tool Use 2026-08-03 arXiv cs.AI (Artificial Intelligence)arXiv — Agents / Tool UsearXiv — AI for SciencearXiv — Evals & Benchmarks 6.5 6.7/6.6/6.3

Agents acting on partial observations face belief-state inference, objective misalignment and planning under uncertainty, and conditioning on full or summarised action-observation histories drags redundant context into every decision. NeSyFS represents the belief state as a knowledge graph and feeds triplets to every module: a fast-thinking module for reactive action, a slow-thinking module doing uncertainty-aware planning structured on twisted sequential Monte Carlo, and a reflection module to correct objective drift.

cs.AI neuro-symbolic planning
#50
Industry 2026-08-01 Hacker News — AI front page 6.5 6.0/6.6/6.8

Reddit shares dropped 23 percent following results attributing a slowdown in user growth to AI assistants answering questions that previously drove search traffic to the platform. The item matters less as a market move than as a datapoint on the referral-traffic mechanism now visible in a public company's reported numbers, since Reddit content is simultaneously a major licensed training and retrieval source.

platforms content market
#51
Efficiency 2026-08-03 arXiv cs.LG (Machine Learning)arXiv — Efficiency (Quantization, MoE, Inference) 6.4 6.5/6.4/6.2

On-policy distillation's rollout cost is dominated by a few long responses delaying batch completion, and existing accelerations cap length with fixed budgets or absolute teacher-student agreement thresholds that do not track learning progress across models or training stages. Adaptive FastOPD expands the horizon only when four teacher-student signals indicate learning near the current boundary has plateaued and the current horizon is sufficiently utilised, making the schedule progress-aware rather than pre-set.

cs.LG distillation rollout
#52
Evaluations & Benchmarks 2026-08-03 arXiv cs.AI (Artificial Intelligence)arXiv — Agents / Tool UsearXiv — Evals & Benchmarks 6.4 6.5/6.4/6.2

Existing agent benchmarks measure static code generation, paper replication or final-answer correctness, none of which test whether an agent can interpret results and act on them. AgentHPOBench provides 30 executable machine-learning tasks across seven research categories, each starting from a validated baseline run, after which the agent performs several sequential interventions while observing accumulated configurations, metrics and logs — turning hyperparameter optimisation into a sequential-decision evaluation rather than a one-shot generation task.

cs.AI benchmark HPO
#53
Agents & Tool Use 2026-08-03 arXiv cs.AI (Artificial Intelligence)arXiv — Agents / Tool UsearXiv — Evals & Benchmarks 6.4 6.5/6.4/6.2

Small language models are attractive for agentic deployment on latency, cost and on-device privacy grounds, but they cannot absorb noisy tool-use supervision the way large models can, making data quality the binding constraint. Data Turnstile takes user-defined API specifications and decomposes multi-turn tool-use interactions into constrained stepwise generation with validation and error-feedback loops, giving explicit control over API diversity, conversation complexity and output correctness. Released open-source.

cs.AI synthetic data tool use SLM
#54
Efficiency 2026-08-03 arXiv cs.LG (Machine Learning)arXiv — Efficiency (Quantization, MoE, Inference)arXiv — Evals & Benchmarks 6.4 6.5/6.4/6.2

Federated fine-tuning of foundation models runs into resource asymmetry: the institutions holding the most valuable domain data cannot host billion-parameter models. Parameter-efficient tuning, pruning and distillation each trade away one of full-model memory reduction, architectural self-containedness, or representational fidelity. FedSLM decomposes with SVD to produce self-contained client models whose low-rank subspaces form nested manifolds, so clients at different capacities occupy compatible subspaces of one server model.

cs.LG federated learning SVD
#55
Research 2026-08-03 arXiv cs.LG (Machine Learning)arXiv — AI for Science 6.4 6.5/6.5/6.1

Solutions occupying larger volumes in parameter space, quantified by Boltzmann entropy, are known to generalise better than those conventional optimisation reaches — the high entropy advantage. This paper asks whether the advantage extends to robustness, meaning retention of learned knowledge when the model is subsequently trained on new information. Using grokking in modular arithmetic as a controlled setting with a noise-injection protocol, it separates apparently-grokked solutions from true equilibrium ones and finds the latter substantially more resistant to forgetting.

cs.LG grokking generalisation
#56
Reinforcement Learning 2026-08-03 arXiv cs.LG (Machine Learning)arXiv — Efficiency (Quantization, MoE, Inference)arXiv — Reinforcement Learning 6.4 6.5/6.4/6.2

NeuroSynth borrows complementary learning systems from biology — rapid hippocampal encoding alongside slower cortical consolidation — and implements it as a dual-pathway continual RL architecture separating rapid task acquisition in a plan pathway from long-term retention in a habit pathway, combined with replay and knowledge distillation. The target is the standard failure where training on new tasks overwrites earlier competence.

cs.LG continual learning RL
#57
Evaluations & Benchmarks 2026-08-03 arXiv cs.AI (Artificial Intelligence)arXiv — Agents / Tool UsearXiv — Evals & Benchmarks 6.4 6.2/6.7/6.2

Financial deployments combine retrieval, proprietary data, tool use, orchestration logic, monitoring and human escalation, yet evaluation stays model-centric — benchmark scores and one-off qualitative reviews treated as evidence of readiness. Drawing on industry experience validating generative AI applications inside financial institutions, the authors argue for validation evidence spanning data, model design, retrieval and generation performance, agent behaviour, governance and implementation as a precondition for production approval.

cs.AI validation finance
#58
Research 2026-08-03 AK (@_akhaliq) Daily PapersarXiv cs.LG (Machine Learning) 6.4 6.5/6.4/6.3

Latent world models depend on the quality of the learned latent distribution, and LeWorldModel regularises toward an isotropic Gaussian with the Epps-Pulley objective — whose corrective gradients, the authors show, vanish rapidly for isolated tail samples, leaving heavy-tailed deviations uncontrolled. QQWorld swaps in a quantile-quantile matching objective that aligns projected latent samples with rank-matched Gaussian quantiles, keeping gradients alive in the tails, plus a cross-batch variant enlarging the ranking pool with detached prior-batch samples. Across four control environments it improves average planning success and yields thinner latent tails.

cs.LG world model regularisation
#59
Efficiency 2026-08-03 arXiv cs.CL (Computation & Language)arXiv — Efficiency (Quantization, MoE, Inference)arXiv — Evals & Benchmarks 6.4 6.5/6.3/6.3

Quantization's memory savings are well characterised; its effect on latency and throughput under realistic, controlled orchestration-level workloads rarely is. This study covers two translation model families across five models from 1.7B to 22B on a single A100 or H100, and finds that pairing a document-chunking strategy with W4A8 or W8A8 quantization improves the latency-throughput Pareto curve across a wide operating range rather than at a single batch size.

cs.CL quantization serving
#60
Infrastructure 2026-07-30 LMSYS Blog (Chatbot Arena) 6.4 6.5/6.4/6.2

RadixArk and Google Cloud announced a partnership to bring SGLang to TPUs so developers can run the same serving stack across hardware backends. Captured from the LMSYS blog listing this run alongside two adjacent posts — Blackwell-native MXFP8 and per-token NVFP4 reinforcement-learning recipes in Miles, and a cleaner quantization stack in SGLang — all published just before this window opened.

SGLang TPU serving
#61
Industry 2026-08-02 Hacker News — AI front page 6.3 6.0/6.3/6.6

A vendor-run audit of 531 stored scans across 458 domains reports that only 10 of 193 sites were ever named across 5,978 buying-intent assistant answers, and that just 1.9 percent of answers named the business being asked about — Gemini 2.9 percent, ChatGPT 1.7, Claude 1.6, Perplexity 1.6. Separately, 8.9 percent of sites block at least one AI crawler, and the blocking is aimed at training rather than retrieval: 38 sites block GPTBot while only 4 block OAI-SearchBot. Caveats are heavy — the corpus is self-selected from a vendor selling AI-visibility services, the headline rests on 193 sites, and the structured-data figures are an acknowledged under-count because the detector missed JSON-LD graph containers until 2 August.

crawlers retrieval measurement
#62
State Space Models 2026-08-03 arXiv cs.CV (Computer Vision)arXiv — State Space ModelsarXiv — Evals & Benchmarks 6.3 6.4/6.3/6.2

Ultra-high-definition restoration has to aggregate spatially recurring degradation cues while preserving local structure, and the usual cost controls — downsampling, window partitioning, cluster-based token reduction — attenuate edges and textures because they discard what shared aggregation represents poorly. CoDe-SSM decouples the two: a global cluster pathway models aggregated context while a separate pathway carries clustering residuals, so fine structure is retained rather than reconstructed.

cs.CV state space models restoration
#63
AI for Science 2026-07-31 Hacker News — AI front pageEmory News 6.3 6.2/6.3/6.3

Automated pose estimation and individual re-identification are letting field researchers run cognitive experiments on wild primate populations at sample sizes that manual video coding could not support, shifting the constraint from analyst hours to camera coverage. The methodological point that generalises is that the bottleneck in observational animal cognition has been annotation throughput rather than experimental design.

ethology computer vision
#64
Evaluations & Benchmarks 2026-08-03 arXiv cs.AI (Artificial Intelligence)arXiv — Agents / Tool UsearXiv — Evals & Benchmarks 6.3 6.4/6.3/6.2

Most agent benchmarks use bounded tasks with immediate success criteria, which cannot detect whether an agent preserves purposeful behaviour across a long horizon. MerchantBench builds a persistent seller-side e-commerce environment where actions constrain future choices, feedback arrives at heterogeneous delays and incoherent behaviour accumulates measurable cost, across product sourcing, listing and pricing control, and cash-flow management.

cs.AI benchmark long horizon
#65
Generative Media 2026-08-03 arXiv cs.CV (Computer Vision)arXiv — Generative Media / DiffusionarXiv — Evals & Benchmarks 6.3 6.4/6.3/6.2

Unifying text-, image- and video-conditioned generation in one model needs a VLM's multimodal understanding wired into a pretrained video diffusion transformer, and existing approaches either inject features from a few manually chosen VLM layers or jointly train architecture-matched understanding and generation streams, which blocks reuse of heterogeneous pretrained backbones. MoRoute learns the routing instead, dynamically selecting which hierarchical VLM representations feed the diffusion transformer per condition.

cs.CV video generation routing
#66
Post-Training 2026-08-03 arXiv cs.CL (Computation & Language)arXiv — Reinforcement LearningarXiv — Evals & Benchmarks 6.3 6.4/6.2/6.2

TAPR reformulates a user prompt into a task-optimised one with the explicit objective of improving downstream performance, trained with GRPO where the reward comes from LLM-as-judge evaluations of both the rewritten prompt and the resulting task output. Reported gains span question answering, summarisation and arithmetic reasoning, with the practical framing that non-expert users should not have to carry prompt engineering themselves.

cs.CL prompting GRPO
#67
Generative Media 2026-08-03 arXiv cs.CV (Computer Vision)arXiv — Generative Media / Diffusion 6.2 6.3/6.3/6.0

Detectors for AI-generated images degrade sharply under post-processing that shifts low-level statistics. RAID exploits the asymmetry between generated and camera-captured images under bit reversal of the input representation, using the difference in detector response before and after the transform as the discriminative signal rather than relying on a single forward pass. Timely given the EU labelling obligations that took effect this window, where watermark absence has to be backstopped by post-hoc detection.

cs.CV detection robustness
#68
Industry 2026-08-02 TechCrunch — AI 6.2 5.9/6.6/6.1

TechCrunch's Equity podcast discusses Sam Altman publicly calling on the industry to pace the rate of AI development, and what that framing means coming from the chief executive of the lab with the largest commercial deployment. The item is commentary on a public statement rather than a policy announcement, and the segment does not report any concrete commitment attached to it.

OpenAI governance
#69
AI for Science 2026-08-03 arXiv cs.LG (Machine Learning)arXiv — AI for Science 6.2 6.3/6.2/6.0

Polymer informatics is usually split across separate models for forward property prediction and inverse structure recommendation, which makes end-to-end design loops brittle. UniPolymer puts prediction, recommendation and evaluation in one framework so candidate structures are scored by the same representation that generated them, closing the loop without a handoff between independently trained components.

cs.LG materials polymers
#70
Industry 2026-08-02 Hacker News — AI front pageThe Wall Street Journal 6.2 5.9/6.3/6.4

Coverage of solo operators reaching seven-figure revenue by delegating functions that previously required staff — support, content, bookkeeping, basic engineering — to AI tooling. The article reached the Hacker News front page largely on the strength of the pattern rather than the sample size; the underlying reporting is anecdotal rather than statistical, and it sits alongside the Apollo wage findings elsewhere in today's digest as the optimistic reading of the same automation trend.

labour automation business
#71
Industry 2026-08-02 Hacker News — AI front page 6.1 5.6/6.2/6.4

A working novelist sets out, step by step, where generative tools would plausibly slot into a professional fiction workflow and why each insertion point is rejected — covering drafting, continuity tracking, research and copy-editing separately rather than treating the question as a single yes or no. It pairs directly with today's market-dilution paper on AI-written self-published fiction.

writing craft
#72
Efficiency 2026-08-03 arXiv cs.LG (Machine Learning)arXiv — Efficiency (Quantization, MoE, Inference) 6.1 6.2/6.1/6.0

GQ-FSL combines federated split learning with quantized activations and gradients, optimising explicitly for energy consumption on edge clients rather than treating energy as a downstream consequence of communication rounds. The framing matters as federated deployments move onto battery-constrained hardware where the binding limit is joules per round, not bandwidth.

cs.LG federated learning quantization
#73
Post-Training 2026-08-03 arXiv cs.CL (Computation & Language)arXiv — AI, Defense & National Security 6.1 6.2/6.1/6.0

InMyStyle is a privacy-first single-user system that adapts small models to rewrite AI-edited text toward an individual's writing style without any instruction prompt at inference. Local helper models build paired training examples from the user's own documents, then LoRA adapters are fine-tuned on bases from 0.5B to 7B parameters, with length-aware generation budgets and automatic chunking. On 219 evaluation pairs from a scientific-paper corpus the composite score plateaus at 0.69 across every model size under both greedy and sampled decoding — the negative result being that capacity is not the bottleneck for this task.

cs.CL LoRA personalisation privacy
#74
Generative Media 2026-08-03 arXiv cs.CV (Computer Vision)arXiv — Generative Media / Diffusion 6.1 6.2/6.2/5.9

Model attribution for generated video normally requires training a classifier per generator, which does not survive new model releases. This retrieval-driven approach matches a query video against a reference bank of known-generator samples in a representation space, so adding a generator means adding references rather than retraining — the practical property that matters when the generator population turns over every few months.

cs.CV attribution video
#75
Frontier LLMs 2026-08-02 Interconnects (Nathan Lambert) 5.9 6.8/7.0/6.9 -1.0 frontier_llm

The twenty-third open-artifacts roundup argues against the consolidation thesis that many observers treated as inevitable as training costs rose by orders of magnitude each year. Laguna S2.1, Inkling and Kimi K3 are used as the counter-evidence: capable open-weight releases continuing to arrive from a widening rather than narrowing set of labs, with enough production utility that the gap to closed frontier models is now argued in terms of specific workloads rather than in general.

open weights analysis
Items
75
Multi-source
61
Long-form (≥7.5)
7
Sources OK / attempted
26 / 119
Top category
Robotic Autonomy
12 items