Augur Dispatch

Chain of evidence

Evidence for 2026-08-22

This frozen page shows Augur's claims and source links for one sent dispatch. Stored spot-checks appear only where the frozen edition supports them; absence is not presented as verification.

As of:

Bundle identity: evidence-bundle-v1-55359ca13af6c043dc3c661e38350b9a79c9d0c5e88d56159eca59fda138e8e0

Format: evidence-bundle-v1 · 46 claims

Assertion 1

Cohere promotes its Transcribe model as beating Whisper Large v3, ElevenLabs Scribe v2, and Qwen3-ASR on the Hugging Face Open ASR Leaderboard, a public ranking of speech-to-text accuracy, and Mistral prices its Voxtral Mini Transcribe V2 at one fifth of Scribe v2 while claiming matched quality Cohere Blog Mistral AI.

Assertion status: No spot-check verdict is published for this assertion.

Voxtral Mini Transcribe V2 costs one-fifth of ElevenLabs' Scribe v2 while matching its quality.

Claim 25418 Label: fact Provenance: primary Recorded

Mistral AI

No stored spot-check names this claim in this edition.

Cohere has released an open-source automatic speech recognition model named Transcribe.

Claim 25676 Label: fact Provenance: primary Recorded

Cohere Blog

No stored spot-check names this claim in this edition.

Cohere Transcribe holds the top rank on Hugging Face's Open ASR Leaderboard with an average word error rate of 5.42%.

Claim 25679 Label: fact Provenance: primary Recorded

Cohere Blog

No stored spot-check names this claim in this edition.

Cohere Transcribe outperforms Whisper Large v3, ElevenLabs Scribe v2, and Qwen3-ASR-1.7B in accuracy on the Hugging Face Open ASR Leaderboard.

Claim 25680 Label: fact Provenance: primary Recorded

Cohere Blog

No stored spot-check names this claim in this edition.

Assertion 2

The clincher came from fresh material: on audio from the same domain but recorded after the models' training cutoffs, the date after which a model saw no new data, the copying habit largely faded, which points at models keying on memorized dataset cues rather than hearing better Hugging Face Blog.

Assertion status: No spot-check verdict is published for this assertion.

An evaluation of 11 widely used open-source automatic speech recognition (ASR) models found that several highest-scoring systems reproduced incorrect transcripts from the VoxPopuli and LibriSpeech datasets even when the audio contradicted them.

Claim 50067 Label: fact Provenance: primary Recorded

Hugging Face Blog

No stored spot-check names this claim in this edition.

In a case study involving VoxPopuli clips with transcription errors, six of the 11 tested models reproduced the benchmark's erroneous reference transcript rather than the audio-faithful transcription.

Claim 50068 Label: fact Provenance: primary Recorded

Hugging Face Blog

No stored spot-check names this claim in this edition.

When models were presented with newly collected audio from the same domain but after their training cutoffs, the benchmark-optimized behavior often weakened or disappeared, suggesting reliance on dataset-specific acoustic cues.

Claim 50069 Label: fact Provenance: primary Recorded

Hugging Face Blog

No stored spot-check names this claim in this edition.

Assertion 3

Hugging Face's own voice evaluation work warned in July that some models may be tuned to reproduce known reference errors or reconstruct masked words absent from the audio, and a Cursor researcher's standing advice is to test models on material released after their training cutoff precisely to separate ability from contamination Hugging Face Blog AI Engineer.

Assertion status: No spot-check verdict is published for this assertion.

Evaluating models on problems released after their training cut-off helps verify that performance drops are due to lack of contamination rather than model capability.

Claim 16686 Label: forecast Provenance: primary Recorded

AI Engineer

No stored spot-check names this claim in this edition.

Transcription word error rates on noise-backed speech were roughly four times higher than on music-backed speech in the evaluation.

Claim 23616 Label: fact Provenance: primary Recorded

Hugging Face Blog

No stored spot-check names this claim in this edition.

Some voice models may be optimized for established public benchmarks by reproducing known errors in reference transcripts or reconstructing masked words not present in the audio.

Claim 23618 Label: fact Provenance: primary Recorded

Hugging Face Blog

No stored spot-check names this claim in this edition.

Voice models often struggle with accents, noise, emotional speech, overlapping speakers, and background noise in real-world conditions.

Claim 23619 Label: fact Provenance: primary Recorded

Hugging Face Blog

No stored spot-check names this claim in this edition.

Assertion 4

A Georgia Tech team traced where the open Olmo model's social reasoning came from using influence functions, a technique that estimates which training documents shaped a given answer, something impossible against a closed vendor's black box Allen Institute for AI.

Assertion status: No spot-check verdict is published for this assertion.

Glenn Matlin, a PhD candidate at Georgia Tech, and co-author Chandreyi Chakraborty used the open-source Olmo 3 model to investigate the origins of its social reasoning capabilities.

Claim 50100 Label: fact Provenance: primary Recorded

Allen Institute for AI

No stored spot-check names this claim in this edition.

The researchers employed influence functions to estimate the impact of individual training documents from the Dolma 3 dataset on the model's answers to specific benchmark questions.

Claim 50101 Label: fact Provenance: primary Recorded

Allen Institute for AI

No stored spot-check names this claim in this edition.

Olmo 3 was selected for the study because it is one of the few large language models with fully public training corpora, checkpoints, and evaluation tools.

Claim 50102 Label: fact Provenance: primary Recorded

Allen Institute for AI

No stored spot-check names this claim in this edition.

Assertion 5

The earlier pairing, DeepSeek V4 Pro backstopped by Claude Fable 5, solved 82.7% of tasks on DeepSWE, a software-engineering benchmark, at $8.28 each, and swapping the backstop to GPT-5.6 Sol already covers 83.0% at just $3.35 per task, which remains the cheapest cascade on the board Together AI Blog Together AI Blog.

Assertion status: No spot-check verdict is published for this assertion.

Running DeepSeek V4 Pro 0813 first and escalating to Claude Fable 5 only when DeepSeek fails solves 82.7% of DeepSWE tasks at a cost of $8.28 per task.

Claim 48029 Label: fact Provenance: primary Recorded

Together AI Blog

No stored spot-check names this claim in this edition.

Claude Fable 5 alone solves 69.7% of DeepSWE tasks at a cost of $21.63 per task.

Claim 48030 Label: fact Provenance: primary Recorded

Together AI Blog

No stored spot-check names this claim in this edition.

DeepSeek V4 Pro 0813 and Claude Fable 5 perform equally at pass@2 (78.5% vs 77.1%) and DeepSeek leads pass@4 (88.5% vs 84.1%) on DeepSWE tasks.

Claim 48032 Label: fact Provenance: primary Recorded

Together AI Blog

No stored spot-check names this claim in this edition.

DeepSeek V4 Pro 0813 costs approximately $0.24 per rollout, 90 times cheaper than Claude Fable 5 at $21.63 per rollout.

Claim 48033 Label: fact Provenance: primary Recorded

Together AI Blog

No stored spot-check names this claim in this edition.

Claude Fable 5 is the most expensive rollout configuration on the DeepSWE leaderboard, while DeepSeek V4 Pro 0813 is among the cheapest.

Claim 48036 Label: fact Provenance: primary Recorded

Together AI Blog

No stored spot-check names this claim in this edition.

Using a cascade approach that runs DeepSeek V4 Pro 0813 first and escalates to GPT-5.6 Sol only if tests fail solves 83.0% of DeepSWE tasks at a cost of $3.35 each, which is 10 points better in accuracy and 60% cheaper than Sol alone.

Claim 48348 Label: fact Provenance: primary Recorded

Together AI Blog

No stored spot-check names this claim in this edition.

GPT-5.6 Sol achieves 72.7% pass@1 success on DeepSWE tasks at $8.37 per rollout, outperforming DeepSeek V4 Pro 0813's 62.8% pass@1 but costs 35 times more per rollout.

Claim 48349 Label: fact Provenance: primary Recorded

Together AI Blog

No stored spot-check names this claim in this edition.

DeepSeek V4 Pro 0813 achieves higher accuracy after multiple attempts, with 88.5% pass@4 versus GPT-5.6 Sol's 85.8%, by leveraging its lower cost that allows for more retries.

Claim 48350 Label: fact Provenance: primary Recorded

Together AI Blog

No stored spot-check names this claim in this edition.

DeepSeek V4 Pro 0813 costs $0.24 per rollout compared to GPT-5.6 Sol's $8.37, making it 35 times cheaper and enabling about 260 solved tasks per $100 versus 9 solved tasks for Sol.

Claim 48351 Label: fact Provenance: primary Recorded

Together AI Blog

No stored spot-check names this claim in this edition.

GPT-5.6 Sol is more likely to produce failures that break already passing tests (20% of failures) compared to DeepSeek V4 Pro 0813 (11% of failures), indicating higher regression risk for Sol.

Claim 48353 Label: fact Provenance: primary Recorded

Together AI Blog

No stored spot-check names this claim in this edition.

GPT-5.6 Sol outperforms DeepSeek V4 Pro 0813 in six of eight task domains on DeepSWE, especially excelling in data modeling and serialization with 92% success compared to Pro, but Pro wins in Rust programming tasks and stateful reactivity.

Claim 48354 Label: fact Provenance: primary Recorded

Together AI Blog

No stored spot-check names this claim in this edition.

DeepSeek V4 Pro 0813 and GPT-5.6 Sol solved 90 of 113 DeepSWE tasks in common, with Pro uniquely solving 10 tasks, Sol uniquely solving 7 tasks, and both failing 6 tasks, together covering 94.7% of tasks.

Claim 48356 Label: fact Provenance: primary Recorded

Together AI Blog

No stored spot-check names this claim in this edition.

The best routing strategy is to run DeepSeek V4 Pro 0813 first and escalate to GPT-5.6 Sol on test failures; this cascading method yields 83.0% task coverage at $3.35 per task, outperforming Sol alone and any one-shot oracle router in accuracy and cost.

Claim 48358 Label: fact Provenance: primary Recorded

Together AI Blog

No stored spot-check names this claim in this edition.

GPT-5.6 Sol is the best single-model choice when first-try correctness and low latency matter, though it costs 35 times more than DeepSeek V4 Pro 0813 and has a higher regression failure rate.

Claim 48359 Label: fact Provenance: primary Recorded

Together AI Blog

No stored spot-check names this claim in this edition.

Assertion 6

The new GLM-5.3-first pairing buys accuracy rather than the lowest bill: it solves 85.9% at $6.61 per task, and GLM-5.3 alone beats Sol on multi-try accuracy, 87.6% to 85.8% Together AI Blog.

Assertion status: No spot-check verdict is published for this assertion.

Running GLM-5.3 first and escalating to GPT-5.6 Sol upon test failure solves 85.9% of DeepSWE tasks at a cost of $6.61 per task.

Claim 50076 Label: fact Provenance: primary Recorded

Together AI Blog

No stored spot-check names this claim in this edition.

GPT-5.6 Sol achieves a 72.7% pass@1 rate on DeepSWE, which is 3.7 percentage points higher than GLM-5.3's 69.0%.

Claim 50077 Label: fact Provenance: primary Recorded

Together AI Blog

No stored spot-check names this claim in this edition.

GLM-5.3 achieves a 87.6% pass@4 rate on DeepSWE, surpassing GPT-5.6 Sol's 85.8%.

Claim 50078 Label: fact Provenance: primary Recorded

Together AI Blog

No stored spot-check names this claim in this edition.

Assertion 7

Nate Jones logged an overnight Codex run costing over $300 against Z.AI's $18-per-month GLM coding plan Nate Jones.

Assertion status: No spot-check verdict is published for this assertion.

The author incurred a cost exceeding $300 for a single overnight Codex run, which they attribute to the tool's continuous validation and repair cycles.

Claim 50144 Label: fact Provenance: primary Recorded

Nate Jones

No stored spot-check names this claim in this edition.

Z.AI's GLM Coding Plan costs $18 per month and is based on GLM-5.3.

Claim 50145 Label: fact Provenance: primary Recorded

Nate Jones

No stored spot-check names this claim in this edition.

API usage for GLM-5.3 costs half as much during off-peak hours, which are weekday afternoons in Singapore.

Claim 50146 Label: fact Provenance: primary Recorded

Nate Jones

No stored spot-check names this claim in this edition.

Assertion 8

Zvi Mowshowitz makes the argument that text watermarking is effectively free and good, and the record he assembles is hard to dismiss: the core technical problem was largely cracked years ago by Scott Aaronson and Hendrik Kirchner during Aaronson's time at OpenAI, the EU's Code of Practice binds the major Western labs that signed it to watermark future models, and Google has been shipping the feature since 2024, most recently in Gemini 3.7 Flash TechCrunch AI Don't Worry About the Vase.

Assertion status: No spot-check verdict is published for this assertion.

Anthropic stated that it will watermark text generated by its AI models, including Claude, to comply with European regulations.

Claim 45157 Label: fact Provenance: primary Recorded

TechCrunch AI

No stored spot-check names this claim in this edition.

For file watermarking, Anthropic is utilizing the C2PA open standard.

Claim 45159 Label: fact Provenance: primary Recorded

TechCrunch AI

No stored spot-check names this claim in this edition.

Scott Aaronson, while working at OpenAI, largely solved AI text watermarking together with Hendrik Kirchner.

Claim 50251 Label: fact Provenance: primary Recorded

Don't Worry About the Vase

No stored spot-check names this claim in this edition.

The European Union Code of Practice, signed by major Western AI labs, requires future AI models to use watermarks.

Claim 50252 Label: fact Provenance: primary Recorded

Don't Worry About the Vase

No stored spot-check names this claim in this edition.

Google implemented AI text watermarking for Gemini 3.7 Flash and has been rolling out the feature since 2024.

Claim 50253 Label: fact Provenance: primary Recorded

Don't Worry About the Vase

No stored spot-check names this claim in this edition.

Assertion 9

The rollout is not friction-free: Anthropic's move has already stirred debate over trust, who gets access to the verifier, and what watermarks do to authorship norms Latent Space.

Assertion status: No spot-check verdict is published for this assertion.

Anthropic’s rollout of Claude text watermarking technology is technically feasible for quality-preserving watermarking but generated significant user trust and policy debates regarding transparency, verifier access, and impacts on authorship norms.

Claim 48249 Label: fact Provenance: primary Recorded

Latent Space

No stored spot-check names this claim in this edition.

Assertion 10

The Department of Justice has reportedly spent close to a year looking at Andreessen Horowitz because Ben Horowitz sits on the board of Databricks while his partner Martin Casado sits on Fivetran's, and those two portfolio companies now compete with each other; the legal hook is reportedly a 112-year-old antitrust statute that almost never gets aimed at venture firms TechCrunch AI.

Assertion status: No spot-check verdict is published for this assertion.

The Department of Justice has reportedly been investigating Andreessen Horowitz for nearly a year regarding board seat conflicts between its partners Ben Horowitz and Martin Casado at portfolio companies Databricks and Fivetran.

Claim 50149 Label: fact Provenance: primary Recorded

TechCrunch AI

No stored spot-check names this claim in this edition.

Andreessen Horowitz partners Ben Horowitz and Martin Casado sit on the boards of Databricks and Fivetran, respectively, which now compete with each other.

Claim 50150 Label: fact Provenance: primary Recorded

TechCrunch AI

No stored spot-check names this claim in this edition.

The DOJ is reportedly utilizing a 112-year-old antitrust law, which is rarely used against venture capital firms, in its investigation of Andreessen Horowitz.

Claim 50151 Label: fact Provenance: primary Recorded

TechCrunch AI

No stored spot-check names this claim in this edition.

Assertion 11

- Hugging Face's contamination findings implicate named test sets, so vendor responses and any re-based leaderboard scores will show who was measuring hearing versus memory. Hugging Face Blog

Assertion status: No spot-check verdict is published for this assertion.

An evaluation of 11 widely used open-source automatic speech recognition (ASR) models found that several highest-scoring systems reproduced incorrect transcripts from the VoxPopuli and LibriSpeech datasets even when the audio contradicted them.

Claim 50067 Label: fact Provenance: primary Recorded

Hugging Face Blog

No stored spot-check names this claim in this edition.

In a case study involving VoxPopuli clips with transcription errors, six of the 11 tested models reproduced the benchmark's erroneous reference transcript rather than the audio-faithful transcription.

Claim 50068 Label: fact Provenance: primary Recorded

Hugging Face Blog

No stored spot-check names this claim in this edition.

When models were presented with newly collected audio from the same domain but after their training cutoffs, the benchmark-optimized behavior often weakened or disappeared, suggesting reliance on dataset-specific acoustic cues.

Claim 50069 Label: fact Provenance: primary Recorded

Hugging Face Blog

No stored spot-check names this claim in this edition.

Assertion 12

- Together AI keeps publishing cascade matchups; the cheap-first advantage is a pattern only if the next pairing repeats it. Together AI Blog

Assertion status: No spot-check verdict is published for this assertion.

Running GLM-5.3 first and escalating to GPT-5.6 Sol upon test failure solves 85.9% of DeepSWE tasks at a cost of $6.61 per task.

Claim 50076 Label: fact Provenance: primary Recorded

Together AI Blog

No stored spot-check names this claim in this edition.

GPT-5.6 Sol achieves a 72.7% pass@1 rate on DeepSWE, which is 3.7 percentage points higher than GLM-5.3's 69.0%.

Claim 50077 Label: fact Provenance: primary Recorded

Together AI Blog

No stored spot-check names this claim in this edition.

GLM-5.3 achieves a 87.6% pass@4 rate on DeepSWE, surpassing GPT-5.6 Sol's 85.8%.

Claim 50078 Label: fact Provenance: primary Recorded

Together AI Blog

No stored spot-check names this claim in this edition.

Assertion 13

- Anthropic's watermark rollout is already drawing complaints from users worried about being caught at work or school, a signal of how enforcement will actually land. TechCrunch AI

Assertion status: No spot-check verdict is published for this assertion.

Anthropic implemented Claude's watermarking policy to comply with the EU AI Act's Transparency Code requirements.

Claim 46008 Label: fact Provenance: primary Recorded

TechCrunch AI

No stored spot-check names this claim in this edition.

The EU AI Act's Transparency Code requires tech companies to label content that is AI-generated or edited in a manner identifiable to computer systems.

Claim 46009 Label: fact Provenance: primary Recorded

TechCrunch AI

No stored spot-check names this claim in this edition.

Assertion 14

- Z.ai launched GLM-5.3 through its API at the same price as GLM-5.2, so today's cascade economics should hold in the near term. Latent Space

Assertion status: No spot-check verdict is published for this assertion.

Z.ai launched GLM-5.3 via API for coding, defensive cyber, and long-horizon agents at the same price as GLM-5.2.

Claim 48969 Label: fact Provenance: primary Recorded

Latent Space

No stored spot-check names this claim in this edition.

Assertion 15

- Georgia Tech's influence-function audit of Olmo suggests more capability-tracing studies on open models are coming, which would extend benchmark skepticism beyond speech. Allen Institute for AI

Assertion status: No spot-check verdict is published for this assertion.

Glenn Matlin, a PhD candidate at Georgia Tech, and co-author Chandreyi Chakraborty used the open-source Olmo 3 model to investigate the origins of its social reasoning capabilities.

Claim 50100 Label: fact Provenance: primary Recorded

Allen Institute for AI

No stored spot-check names this claim in this edition.

The researchers employed influence functions to estimate the impact of individual training documents from the Dolma 3 dataset on the model's answers to specific benchmark questions.

Claim 50101 Label: fact Provenance: primary Recorded

Allen Institute for AI

No stored spot-check names this claim in this edition.

Olmo 3 was selected for the study because it is one of the few large language models with fully public training corpora, checkpoints, and evaluation tools.

Claim 50102 Label: fact Provenance: primary Recorded

Allen Institute for AI

No stored spot-check names this claim in this edition.