Podcast · Every explainer, listenable
Xiaohu · Best Blog Podcast
The latest in AI, explained in plain words — a two-host dialogue show.
85 episodesTwo hostsDaily🔒 49 member-only

Two API settings, not a new model: OpenAI reports GPT-5.6 Sol jumping from 13.3% to 38.3% on ARC-AGI-3
The model stayed the same; the harness did not—retained private reasoning plus compaction instead of deletion cut output tokens per game to about one-sixth.

How GitHub’s official harness guide becomes a clear, repeatable 8-step workflow
He even installs two skills in the piece itself — and that “mostly” at the end of the original title is doing real work.

OpenAI's gpt-transcribe and gpt-live-transcribe halve real-world word error rate, cut price 25%
The real story is in the control test: three older models fed the same background context showed no gains — and two actually got worse.

OpenAI analyzed 800,000 work-related messages: most of what people use AI for is not what their job was supposed to be
The exact same dataset produces task crossover rates of 43.5% and 65%–82%, differing only in how the denominator is sliced.

Kimi K3 technical report: How three architectural changes boosted compute efficiency 2.5x
Published 11 days after launch, the 47-page paper focuses on efficiency, delivering 2.5 times the performance of K2 on the same compute budget.

Terence Tao at the International Congress of Mathematicians: AI will generate more proofs, but mathematics won't speed up
Mathematical research is a five-step pipeline, and Terence Tao argues AI only speeds up step one—the remaining four steps grow progressively slower and rely even more heavily on human judgment.

The story behind Claude Design and 10 practical tips from its creator
The three most counterintuitive takeaways: ask for ten variations at once, stick to wireframes when details do not matter, and handle the final mile yourself.

Anthropic's official guide to prompting Opus 5: subtract, don't add
Eight prompts you can copy straight into your workflow — plus one counterintuitive lesson: telling a review prompt to flag only high-severity issues can genuinely make it report less.

Anthropic slashes Claude Code's system prompt by 80%: six new rules for writing context
Rules gave way to judgment calls and examples gave way to interfaces — one tool description shrank from roughly 9,100 characters to a single sentence and an enum.

Why US unemployment hasn't risen even as AI races ahead, according to Anthropic's economics chief
His take: AI is amplifying workers rather than replacing them—killing tasks is not the same as killing jobs.

YC names the startup ideas it's eager to fund right now
From AI tutors for kids to data centers built at sea — and, for the first time, a request from the sitting US Secretary of the Army.

OpenAI test model escapes its sandbox, breaches Hugging Face's systems
To probe how far the model's attack skills could go, researchers dialed down its refusal to engage in cyberattacks — and it went from a sandbox meant only for installing packages to breaking into another company's production database.

Chroma founder's 12 rules for the AI-era company: teach AI the job first, hire only if it can't learn
The bottleneck at a company, he argues, has shifted from how fast people work to how good their taste and judgment are — a personal take, not a data-backed study.

Study finds a sweet spot for AI-assisted writing — not too much, not too little
Three separate trials of the same study all landed on the middle ground — and the columnist who tried ChatGPT on a movie synopsis says he'd still rather write it himself.

Have Chinese AI models caught up to the US frontier? A look at the real gap
Kimi K3's launch set off claims that China has caught up with the US, but a different measuring stick puts the lag at more than triple — and even flips which side is accelerating faster.

An AI agent breached Hugging Face's internal network, then commercial LLMs mistook the investigators for attackers
The breach ran all weekend and left over 17,000 action logs behind — the team only made sense of it after turning to GLM 5.2, an open-source model running in a self-hosted environment.

The model you picked through a router might be swapped — one random number proves it
The fingerprint distance between two samples of the same model has a median of 0.140. One API marketed as a proprietary in-house flagship scores 0.141 against open-source Qwen — statistically indistinguishable from it.

Google engineer turns a year of harness engineering into 12 rules and a copy-paste playbook
The counterintuitive core: lock the model and agent in place, don't touch a single parameter, and change only the context and tools around them.

SenseTime launches SenseNova U1 Pro: native 8K images that stay sharp even when you zoom into dense small text
At WAIC, the model showed off a 4:1 ink-wash scroll and a 22-panel storyboard sequence. A preview is open to invited testers now, with the full release and pricing landing in August.

Inside Anthropic: the harness wrapped around agents is coming apart
Models are now capable enough to plan their own steps, turning the orchestration built around them into a straitjacket — a 16-minute internal conversation on what a thinner harness looks like.

Cerebras details how it built a knowledge base that leaves data where it lives and lets AI fetch it
Three months in, it's fielding more than 15,000 queries a day — and Cerebras published the actual parameters behind its four-way scoring, thread distillation, and rank fusion.

a16z: the software selloff is only killing companies without a moat
A16z's weekly chart deck also pushes back on three claims — that cheap models are undercutting frontier labs, that AI is stealing jobs, and that data centers are driving up electricity prices — with the data mostly telling the opposite story.

AI can copy your work in seconds — but not these eight things
Kevin Kelly wrote "Better Than Free" back in 2008, and the AI era just proved him right: once copies are free, what sells is whatever can't be copied.

Google's gemma-trainer skill fine-tunes a custom Gemma right on your own machine
This open-source training playbook locks down all three fine-tuning paths — supervised fine-tuning, preference alignment, and reward scoring — plus the LoRA parameters, pairs with Unsloth, and runs on a consumer GPU with just 8GB of VRAM.

The CTO is an AI too: how a nearly all-AI team shipped a product feature
Of the team's 15 members, only the founder is human — the rest are AI. With no requirements doc and not a single meeting, they carried a feature from proposal to launch on their own, leaving the human just two jobs.

A single wrapper pushes Opus 4.8 and Fable 5 to 99% on ARC-AGI-3
A harness called Schema has models turn each game's rules into a runnable, verified program before making a move. Across 25 public rounds it self-reported 98.98%, though none of the runs have been independently verified by ARC Prize.

LM Studio launches Bionic, a Codex-style agent for open models
Run it locally for free with your data staying on-device, or switch to the cloud for more horsepower — zero data retention by default. The preview is live today on Mac and Windows.

Moonshot AI ships Kimi K3, the world's first 3-trillion-parameter open model
2.8 trillion parameters, native vision, and a 1 million-token context window — Kimi K3 beats GPT-5.6 Sol on several agent benchmarks, with the app and API live today

How Should Websites Optimize in the AI Search Era? Google's Official SEO Guide
In its first systematic statement on optimizing for AI Overviews and AI Mode, Google Search also called out a batch of AEO/GEO buzzwords by name — and told sites to drop them.

You just hired a million bad employees
Companies are handing every employee unlimited AI agents and token budgets — and bad workflows now replicate by the second. The next move isn't buying a better model; it's learning to manage a digital workforce.

How much did this order really make? Retail finance turns to AI
Customers buy online, pick up in store, and return across locations—one purchase, many ledgers. Genie is framed as a finance sidekick for true profit, trapped cash, and fewer markdowns.

PrismML Stuffs a 27B Model into Your iPhone — With Barely Any IQ Drop
It squeezes a ~54GB 27B model down to about 3.9–5.9GB: it runs on-device, and average scores still hold about nine-tenths. Core value first; technical detail and the fine print come later.

DeepMind's Demis Hassabis: give frontier models a 30-day safety check before release — or keep them off the market
Hassabis wants the US to stand up a FINRA-style body that defines frontier models with a moving benchmark — a voluntary protocol now, a hard gate to market later.

Princeton professor at ICML on AI and work: how should individuals adapt?
In ~24 months, model capability rose sharply; SAGE’s composite reliability metric rose only 5–10 percentage points

Anthropic's playbook for AI-native startups: four stages, graduation criteria, and ready-to-copy prompts
Anthropic's 36-page playbook breaks down the graduation bar and common pitfalls for four startup stages, complete with matching Claude prompts you can copy straight in.

DoorDash's AI Shopping Assistant: It's Not the Model, It's the Tools and Memory
After the memory system launched, grocery checkout conversion rose about 24%, and automated evals scaled daily test volume from 1 human-reviewed case to 2,000+.

How Microsoft Ships AI Agents at Enterprise Scale: Breaking a Production Agent into 5 Layers
Retrieval becomes a sub-agent that plans and retries; evaluation upgrades from "does it run" to "did it do it right"

Anthropic Analyzed 300,000 Conversations: Claude's Values Shift by Language
Across 3 models and 20 languages: English is the most cautious and in-depth, Russian the most exacting, Hindi the warmest, and Chinese sits closest to the global average

Sakana AI's Brainless 3D Cell Bricks Self-Organize With Biology-Like Swarm Intelligence
With no central brain in charge, nearly 200 simple smart cubes figure out what shape they've formed just by talking to their neighbors—and can even sense where to "regrow" after damage. The self-recognition part already works on physical bricks; damage localization and regeneration still happen mostly in simulation.

Microsoft CEO Satya Nadella: Companies Face a ‘Reverse Information Paradox,’ Paying for AI While Handing Over Their Internal Knowledge and Experience
Companies pay for intelligence twice: once in model fees, and again in the proprietary know-how required to make the model genuinely useful.

Meta AI's Proactive Memory Agent Teaches Models When Not to Remind You, Boosting Terminal-Bench Accuracy by 8.3 Points
The paper claims the code is open-sourced — but the repo turns out to be empty, without a single commit ever pushed.

Ploy Switched Production AI from Opus 4.8 to GPT-5.6 Sol: Cutting Latency by Over Half and Costs by 27% Without Losing Quality
Swapping models isn't just swapping an API: eval frameworks, tool parameters, caching, and reasoning traces — four invisible pitfalls, unpacked and fixed one by one

Claude Cowork's Top Use Case Is Office Admin (33.4%), Not Coding (8.7%): Interface Shapes How AI Gets Used
Based on 1.2M+ conversations across 600,000+ organizations: content creation ranks second at 16.4%, together accounting for nearly half of all usage.

Low-Margin Industries Are AI's Biggest Winners
Manufacturing, logistics, warehousing, and labor services have been stuck at single-digit margins for years — cutting coordination costs alone can multiply their profits.

The Package AI Told You to Install Doesn't Exist — Hackers Already Named Their Malware After It
576K samples, 16 models tested: 19.7% of AI-recommended packages are hallucinations, and 43% keep generating the same fake name.

OpenAI Staff Reveal How They Really Use GPT-5.6: Unlimited Quota, Not Always Maxed Out
The team dodged questions about benchmark gaming, faced backlash from longtime users over the desktop app merger, and admitted they're "still figuring it out."

Mira Murati, OpenAI's Former CTO: AI Should Amplify Humans, Not Replace Them
From tacit knowledge to interaction bandwidth to model alignment — why AI's progress still can't do without humans.

One AI Per Department Backfired — Sierra's Single Assistant Now Handles 70% of Its Code
A retrospective: since launch, Pinecone has run 75,000 sessions, served 600+ employees, and connected 37 internal systems via MCP Gateway.

OpenAI's Official Guide: 9 Copy-Ready Codex Prompt Workflows for Fixing Bugs, Turning Screenshots into Prototypes, and Cloud Refactors
OpenAI consolidates prompting tips scattered across its product pages into one framework — goal, context, output, constraints — plus dedicated Codex workflow examples.

What Does Each Brain Region Like to Watch? EPFL Evolved AI-Generated Videos to Find Out
All results are computer simulation predictions from a brain "digital twin" model, not yet validated with real human brain imaging.

LangChain Ships OpenWiki 0.1.0: Proactive Memory for AI Agents, Auto-Syncing Gmail, Notion, Git, and X into a Local Wiki
No manual context-feeding required — the local Markdown wiki refreshes on a set schedule; a Slack connector is coming soon.

Google Cloud Puts a Proxy Model Inside AlloyDB: In-Database AI Inference, Up to 23,000x Faster
Local small models replace cloud AI calls — the numbers come from Google's internal testing; proxy models are currently limited to the ai.if function and still in preview

Google Research Unveils SensorFM: Trained on a Trillion Minutes of Wearable Data, It Wins 33 of 35 Health Tasks
Pretrained on 2 billion hours of wearable data from 5 million people, a frozen encoder with just a linear head beats supervised baselines on 34 of 35 health tasks.

Every's 9-Person Team Ran GPT-5.6 Sol for a Month: Here's the Verdict
First-hand impressions from four scenarios: coding, writing, knowledge work, and agents

Claude Code's Two Knobs, Explained: One for Capability, One for Effort
A ClaudeDevs deep dive untangles two knobs that both seem to promise a better answer: switching models swaps in a different frozen set of weights, while dialing up effort changes how willing it is to read more files, run more tests, and double-check before handing in the result.

Gemma 4 Technical Report: How a Small Open Model Takes On Large Ones with Reasoning and Memory Efficiency
Not a feature list — three engineering tracks turned at once: reasoning lifts intelligence, an efficiency stack cuts cost, native multimodality expands input.

LangChain Tunes the Harness, Not the Model — Nemotron 3 Ultra Closes In on Opus 4.8 at 1/10 the Cost
Three levers — system prompt, tool descriptions, middleware — push the Deep Agents suite from a typical ~0.80 to 0.84, topping out at 0.86 against Opus's 0.87.

AI Self-Improvement Starts Outside the Model: Lilian Weng on Harness Engineering
Former OpenAI safety lead surveys nearly 30 papers: from prompt tweaks to self-modifying code, DGM pushed coding ability from 20% to 50%.

HKUDS Open-Sources AI Tutor DeepTutor — 20K GitHub Stars in 111 Days
Paper shows a 10.8% gain in personalization and 29.4% in reasoning, fully open-source under Apache 2.0

Anthropic's Playbook: Fable 5 Advises, Sonnet 5 Foots the Bill
In advisor mode Fable 5 just gives advice; in orchestrator mode it delegates tasks — either way, the cheaper Sonnet 5 ends up doing most of the work

Liquid AI's Antidoom Fixes Reasoning Models' Doom Loops With a Single Token — Now Open-Sourced
By fine-tuning only the single token where the doom loop begins, both models' loop rates drop to around 1%

Cloudflare's New Gateway Lets AI Pay You Automatically to Crawl Your Site
Expanding from charging only AI crawlers to charging any caller, now in early access waitlist

Anthropic Discovers a Brain-Like 'Inner Workspace' in Claude — It Evolved, Wasn't Designed
It makes up less than 10% of the model — remove it and Claude can still talk, but its reasoning collapses to zero. Anthropic is already using it to catch fabricated data and spot when Claude senses it's being tested.

Dartmouth Put AI Grading to the Test: Students Called It Rigid — But Users Scored Higher
In a 151-student trial, short-answer questions moved scores more than multiple choice, while almost no one touched the AI help sidebar

An AI Engineer Built His Nonverbal Autistic Son a Communication App — Speech Jumped 5x, and It Became a Real Business
No funding, no team — built purely to solve his own kid's problem. Clinics and schools started asking to use it anyway.

AIEWF closing debate pits agentic coding-loop hype against engineering discipline
A live audience vote couldn't be tallied because the venue lights were too bright to count hands, but a companion survey found 95% of teams already use agents while 59% worry about mounting technical debt.

Wealthy American Families Ditch Public School for $75,000-a-Year AI Private Schools
While traditional schools are still figuring out AI, Silicon Valley and Wall Street families are already voting with their wallets.

AI Is Torching Junior Dev Jobs: Coding Is Becoming a Basic Skill, Not a Career
US developer employment among 22-to-25-year-olds has fallen 19% in three years, even as new GitHub sign-ups hit their fastest growth ever.

NVIDIA Research Unveils HORIZON: Unattended Agents Push Full RTL Chip Design Benchmark to 100% Pass Rate
The paper is the first to run an agent through an entire RTL benchmark suite fully unattended—most tasks clear in two or three rounds, but the hardest one takes 82 iterations

Anthropic Analyzed 400K Claude Code Sessions: Expertise Beats Coding Skill
A 7-month analysis of sessions from 235,000 users: verified experts succeed at nearly double the rate of novices — yet the top 10 professions differ by no more than 7 percentage points.

Dan Koe's Guide to Persuasion: 3 Tensions, 5 Levers — Steal It for Your Copy
Create psychological tension first, then offer a first step too small to refuse — it works for writing, selling, and job hunting alike.

Sakana AI Co-Founder: The Ability to Swap Models on Demand Is Itself a Deterrent
On June 12, U.S. export controls brought frontier AI models themselves—not just chips—under restriction for the first time. His answer: master multi-model orchestration.

Working With Claude Fable 5: The Real Skill Is Finding Your Own Unknowns
Anthropic's Thariq argues the quality of your work with Claude Fable 5 hinges on how clearly you can name your own unknowns. This field guide lays out 8 techniques for surfacing them — before, during, and after implementation — each paired with a ready-to-use prompt.

As AI Gets Cheaper, Renting Generic Intelligence Gets Riskier
Investor Chamath Palihapitiya: intelligence is getting cheap like phones, and expert judgment is now available to everyone — the real moat is encoding your proprietary experience into your own system, not renting the same generic AI as your competitors

Every's Head of Consulting Can't Code — She Still Builds Her Own Work Systems With Codex
She used the same pattern to build an email triage tool and a caregiving app for her dad — the method is repeatable, though the full prompts remain unpublished.

Same AI Model, Opposite Outcomes: Why Some Companies Compound Gains While Others Get Nothing
Four implementation principles, backed by real-world data from L'Oréal, Lyft, and Rakuten

Alibaba Open-Sources Page Agent: Agents Embedded in the Webpage, Reading Text Instead of Screenshots
MIT-licensed and model-agnostic — works with any OpenAI-compatible text model, though for now it only handles a single page view

The AI Leverage Ladder for PMs: Two Paths to Higher Output, Plus 3 Copy-Paste Prompts
An instructor who has trained 30,000+ PMs breaks down two AI leverage ladders — from copy-paste to end-to-end delivery, and from web prototypes to production PRs.

Claude Code's Official Playbook: 4 Levels of Agent Loops to Unattended
From manual confirmation to fully unattended, the Claude Code team lays out a 4-level loop taxonomy with practical guidance

Meituan Launches LongCat-2.0: 1.6-Trillion-Parameter Model Trained Entirely on Domestic Chips, No Nvidia GPUs
Trained on 50,000+ domestic AI chips and 35 trillion tokens; most benchmarks come from Meituan's own evaluation framework, and the weights aren't truly open for download yet

Anthropic Launches Claude Science: An AI Workbench for Scientists with 60+ Built-In Research Skills
Now in open beta. A coordinator agent marshals a team of expert agents to do the work, with a reviewer agent at the end dedicated to catching errors in citations and numbers — compute gets outsourced to AI, but raw data never leaves your local machine.

Bridgewater Built a Financial-Filtering Model with 84.7% Accuracy — and Open-Sourced the Method
Partnering with Thinking Machines, they fine-tuned an open-source model on expert-labeled data: 29.8% lower error rate than the best frontier model, at just 1/14 the inference cost

AI Replaced 700 Agents at Klarna. Then Quality Dropped.
Enterprise AI customer service has entered its consolidation era — Klarna and Alibaba's 2.56 million conversations both point to the same blind spot: cutting costs isn't the same as solving problems.
—
0:000:00
