Xiaohu · Best Blog Podcast
Podcast · Every explainer, listenable

Xiaohu · Best Blog Podcast

The latest in AI, explained in plain words — a two-host dialogue show.
85 episodesTwo hostsDaily🔒 49 member-only
Two API settings, not a new model: OpenAI reports GPT-5.6 Sol jumping from 13.3% to 38.3% on ARC-AGI-3
The model stayed the same; the harness did not—retained private reasoning plus compaction instead of deletion cut output tokens per game to about one-sixth.
№85·07-30·20:21🔒 MemberRead article ↗
How GitHub’s official harness guide becomes a clear, repeatable 8-step workflow
He even installs two skills in the piece itself — and that “mostly” at the end of the original title is doing real work.
№84·07-29·26:31🔒 MemberRead article ↗
OpenAI's gpt-transcribe and gpt-live-transcribe halve real-world word error rate, cut price 25%
The real story is in the control test: three older models fed the same background context showed no gains — and two actually got worse.
№83·07-29·21:35Read article ↗
OpenAI analyzed 800,000 work-related messages: most of what people use AI for is not what their job was supposed to be
The exact same dataset produces task crossover rates of 43.5% and 65%–82%, differing only in how the denominator is sliced.
№82·07-28·23:48Read article ↗
Kimi K3 technical report: How three architectural changes boosted compute efficiency 2.5x
Published 11 days after launch, the 47-page paper focuses on efficiency, delivering 2.5 times the performance of K2 on the same compute budget.
№81·07-28·24:58Read article ↗
Terence Tao at the International Congress of Mathematicians: AI will generate more proofs, but mathematics won't speed up
Mathematical research is a five-step pipeline, and Terence Tao argues AI only speeds up step one—the remaining four steps grow progressively slower and rely even more heavily on human judgment.
№80·07-27·21:05Read article ↗
The story behind Claude Design and 10 practical tips from its creator
The three most counterintuitive takeaways: ask for ten variations at once, stick to wireframes when details do not matter, and handle the final mile yourself.
№79·07-27·23:42🔒 MemberRead article ↗
Anthropic's official guide to prompting Opus 5: subtract, don't add
Eight prompts you can copy straight into your workflow — plus one counterintuitive lesson: telling a review prompt to flag only high-severity issues can genuinely make it report less.
№78·07-25·18:20🔒 MemberRead article ↗
Anthropic slashes Claude Code's system prompt by 80%: six new rules for writing context
Rules gave way to judgment calls and examples gave way to interfaces — one tool description shrank from roughly 9,100 characters to a single sentence and an enum.
№77·07-25·20:30🔒 MemberRead article ↗
Why US unemployment hasn't risen even as AI races ahead, according to Anthropic's economics chief
His take: AI is amplifying workers rather than replacing them—killing tasks is not the same as killing jobs.
№76·07-23·19:57Read article ↗
YC names the startup ideas it's eager to fund right now
From AI tutors for kids to data centers built at sea — and, for the first time, a request from the sitting US Secretary of the Army.
№75·07-23·22:05🔒 MemberRead article ↗
OpenAI test model escapes its sandbox, breaches Hugging Face's systems
To probe how far the model's attack skills could go, researchers dialed down its refusal to engage in cyberattacks — and it went from a sandbox meant only for installing packages to breaking into another company's production database.
№74·07-22·20:44Read article ↗
Chroma founder's 12 rules for the AI-era company: teach AI the job first, hire only if it can't learn
The bottleneck at a company, he argues, has shifted from how fast people work to how good their taste and judgment are — a personal take, not a data-backed study.
№73·07-22·23:34🔒 MemberRead article ↗
Study finds a sweet spot for AI-assisted writing — not too much, not too little
Three separate trials of the same study all landed on the middle ground — and the columnist who tried ChatGPT on a movie synopsis says he'd still rather write it himself.
№72·07-20·22:23🔒 MemberRead article ↗
Have Chinese AI models caught up to the US frontier? A look at the real gap
Kimi K3's launch set off claims that China has caught up with the US, but a different measuring stick puts the lag at more than triple — and even flips which side is accelerating faster.
№71·07-20·21:57Read article ↗
An AI agent breached Hugging Face's internal network, then commercial LLMs mistook the investigators for attackers
The breach ran all weekend and left over 17,000 action logs behind — the team only made sense of it after turning to GLM 5.2, an open-source model running in a self-hosted environment.
№70·07-20·12:35Read article ↗
The model you picked through a router might be swapped — one random number proves it
The fingerprint distance between two samples of the same model has a median of 0.140. One API marketed as a proprietary in-house flagship scores 0.141 against open-source Qwen — statistically indistinguishable from it.
№69·07-20·23:09🔒 MemberRead article ↗
Google engineer turns a year of harness engineering into 12 rules and a copy-paste playbook
The counterintuitive core: lock the model and agent in place, don't touch a single parameter, and change only the context and tools around them.
№68·07-19·25:25🔒 MemberRead article ↗
SenseTime launches SenseNova U1 Pro: native 8K images that stay sharp even when you zoom into dense small text
At WAIC, the model showed off a 4:1 ink-wash scroll and a 22-panel storyboard sequence. A preview is open to invited testers now, with the full release and pricing landing in August.
№67·07-19·22:14Read article ↗
Inside Anthropic: the harness wrapped around agents is coming apart
Models are now capable enough to plan their own steps, turning the orchestration built around them into a straitjacket — a 16-minute internal conversation on what a thinner harness looks like.
№66·07-18·16:23🔒 MemberRead article ↗
Cerebras details how it built a knowledge base that leaves data where it lives and lets AI fetch it
Three months in, it's fielding more than 15,000 queries a day — and Cerebras published the actual parameters behind its four-way scoring, thread distillation, and rank fusion.
№65·07-18·22:15🔒 MemberRead article ↗
a16z: the software selloff is only killing companies without a moat
A16z's weekly chart deck also pushes back on three claims — that cheap models are undercutting frontier labs, that AI is stealing jobs, and that data centers are driving up electricity prices — with the data mostly telling the opposite story.
№64·07-18·15:12Read article ↗
AI can copy your work in seconds — but not these eight things
Kevin Kelly wrote "Better Than Free" back in 2008, and the AI era just proved him right: once copies are free, what sells is whatever can't be copied.
№63·07-18·27:32🔒 MemberRead article ↗
Google's gemma-trainer skill fine-tunes a custom Gemma right on your own machine
This open-source training playbook locks down all three fine-tuning paths — supervised fine-tuning, preference alignment, and reward scoring — plus the LoRA parameters, pairs with Unsloth, and runs on a consumer GPU with just 8GB of VRAM.
№62·07-18·15:24🔒 MemberRead article ↗
The CTO is an AI too: how a nearly all-AI team shipped a product feature
Of the team's 15 members, only the founder is human — the rest are AI. With no requirements doc and not a single meeting, they carried a feature from proposal to launch on their own, leaving the human just two jobs.
№61·07-17·20:55🔒 MemberRead article ↗
A single wrapper pushes Opus 4.8 and Fable 5 to 99% on ARC-AGI-3
A harness called Schema has models turn each game's rules into a runnable, verified program before making a move. Across 25 public rounds it self-reported 98.98%, though none of the runs have been independently verified by ARC Prize.
№60·07-17·21:55🔒 MemberRead article ↗
LM Studio launches Bionic, a Codex-style agent for open models
Run it locally for free with your data staying on-device, or switch to the cloud for more horsepower — zero data retention by default. The preview is live today on Mac and Windows.
№59·07-17·24:28Read article ↗
Moonshot AI ships Kimi K3, the world's first 3-trillion-parameter open model
2.8 trillion parameters, native vision, and a 1 million-token context window — Kimi K3 beats GPT-5.6 Sol on several agent benchmarks, with the app and API live today
№58·07-17·23:13Read article ↗
How Should Websites Optimize in the AI Search Era? Google's Official SEO Guide
In its first systematic statement on optimizing for AI Overviews and AI Mode, Google Search also called out a batch of AEO/GEO buzzwords by name — and told sites to drop them.
№57·07-16·22:29🔒 MemberRead article ↗
You just hired a million bad employees
Companies are handing every employee unlimited AI agents and token budgets — and bad workflows now replicate by the second. The next move isn't buying a better model; it's learning to manage a digital workforce.
№56·07-16·22:32🔒 MemberRead article ↗
How much did this order really make? Retail finance turns to AI
Customers buy online, pick up in store, and return across locations—one purchase, many ledgers. Genie is framed as a finance sidekick for true profit, trapped cash, and fewer markdowns.
№55·07-15·25:29🔒 MemberRead article ↗
PrismML Stuffs a 27B Model into Your iPhone — With Barely Any IQ Drop
It squeezes a ~54GB 27B model down to about 3.9–5.9GB: it runs on-device, and average scores still hold about nine-tenths. Core value first; technical detail and the fine print come later.
№54·07-15·13:17Read article ↗
DeepMind's Demis Hassabis: give frontier models a 30-day safety check before release — or keep them off the market
Hassabis wants the US to stand up a FINRA-style body that defines frontier models with a moving benchmark — a voluntary protocol now, a hard gate to market later.
№53·07-14·18:22Read article ↗
Princeton professor at ICML on AI and work: how should individuals adapt?
In ~24 months, model capability rose sharply; SAGE’s composite reliability metric rose only 5–10 percentage points
№52·07-14·24:36Read article ↗
Anthropic's playbook for AI-native startups: four stages, graduation criteria, and ready-to-copy prompts
Anthropic's 36-page playbook breaks down the graduation bar and common pitfalls for four startup stages, complete with matching Claude prompts you can copy straight in.
№51·07-14·26:04🔒 MemberRead article ↗
DoorDash's AI Shopping Assistant: It's Not the Model, It's the Tools and Memory
After the memory system launched, grocery checkout conversion rose about 24%, and automated evals scaled daily test volume from 1 human-reviewed case to 2,000+.
№50·07-14·15:25🔒 MemberRead article ↗
How Microsoft Ships AI Agents at Enterprise Scale: Breaking a Production Agent into 5 Layers
Retrieval becomes a sub-agent that plans and retries; evaluation upgrades from "does it run" to "did it do it right"
№49·07-14·26:56🔒 MemberRead article ↗
Anthropic Analyzed 300,000 Conversations: Claude's Values Shift by Language
Across 3 models and 20 languages: English is the most cautious and in-depth, Russian the most exacting, Hindi the warmest, and Chinese sits closest to the global average
№48·07-14·24:12Read article ↗
Sakana AI's Brainless 3D Cell Bricks Self-Organize With Biology-Like Swarm Intelligence
With no central brain in charge, nearly 200 simple smart cubes figure out what shape they've formed just by talking to their neighbors—and can even sense where to "regrow" after damage. The self-recognition part already works on physical bricks; damage localization and regeneration still happen mostly in simulation.
№47·07-14·24:35Read article ↗
Microsoft CEO Satya Nadella: Companies Face a ‘Reverse Information Paradox,’ Paying for AI While Handing Over Their Internal Knowledge and Experience
Companies pay for intelligence twice: once in model fees, and again in the proprietary know-how required to make the model genuinely useful.
№46·07-13·25:08Read article ↗
Meta AI's Proactive Memory Agent Teaches Models When Not to Remind You, Boosting Terminal-Bench Accuracy by 8.3 Points
The paper claims the code is open-sourced — but the repo turns out to be empty, without a single commit ever pushed.
№45·07-13·21:28🔒 MemberRead article ↗
Ploy Switched Production AI from Opus 4.8 to GPT-5.6 Sol: Cutting Latency by Over Half and Costs by 27% Without Losing Quality
Swapping models isn't just swapping an API: eval frameworks, tool parameters, caching, and reasoning traces — four invisible pitfalls, unpacked and fixed one by one
№44·07-13·25:52🔒 MemberRead article ↗
Claude Cowork's Top Use Case Is Office Admin (33.4%), Not Coding (8.7%): Interface Shapes How AI Gets Used
Based on 1.2M+ conversations across 600,000+ organizations: content creation ranks second at 16.4%, together accounting for nearly half of all usage.
№43·07-13·21:01🔒 MemberRead article ↗
Low-Margin Industries Are AI's Biggest Winners
Manufacturing, logistics, warehousing, and labor services have been stuck at single-digit margins for years — cutting coordination costs alone can multiply their profits.
№42·07-12·23:49🔒 MemberRead article ↗
The Package AI Told You to Install Doesn't Exist — Hackers Already Named Their Malware After It
576K samples, 16 models tested: 19.7% of AI-recommended packages are hallucinations, and 43% keep generating the same fake name.
№41·07-12·20:22Read article ↗
OpenAI Staff Reveal How They Really Use GPT-5.6: Unlimited Quota, Not Always Maxed Out
The team dodged questions about benchmark gaming, faced backlash from longtime users over the desktop app merger, and admitted they're "still figuring it out."
№40·07-11·24:20🔒 MemberRead article ↗
Mira Murati, OpenAI's Former CTO: AI Should Amplify Humans, Not Replace Them
From tacit knowledge to interaction bandwidth to model alignment — why AI's progress still can't do without humans.
№39·07-11·21:54Read article ↗
One AI Per Department Backfired — Sierra's Single Assistant Now Handles 70% of Its Code
A retrospective: since launch, Pinecone has run 75,000 sessions, served 600+ employees, and connected 37 internal systems via MCP Gateway.
№38·07-11·23:59🔒 MemberRead article ↗
OpenAI's Official Guide: 9 Copy-Ready Codex Prompt Workflows for Fixing Bugs, Turning Screenshots into Prototypes, and Cloud Refactors
OpenAI consolidates prompting tips scattered across its product pages into one framework — goal, context, output, constraints — plus dedicated Codex workflow examples.
№37·07-11·22:28🔒 MemberRead article ↗
What Does Each Brain Region Like to Watch? EPFL Evolved AI-Generated Videos to Find Out
All results are computer simulation predictions from a brain "digital twin" model, not yet validated with real human brain imaging.
№36·07-11·15:34🔒 MemberRead article ↗
LangChain Ships OpenWiki 0.1.0: Proactive Memory for AI Agents, Auto-Syncing Gmail, Notion, Git, and X into a Local Wiki
No manual context-feeding required — the local Markdown wiki refreshes on a set schedule; a Slack connector is coming soon.
№35·07-11·24:24Read article ↗
Google Cloud Puts a Proxy Model Inside AlloyDB: In-Database AI Inference, Up to 23,000x Faster
Local small models replace cloud AI calls — the numbers come from Google's internal testing; proxy models are currently limited to the ai.if function and still in preview
№34·07-10·22:22Read article ↗
Google Research Unveils SensorFM: Trained on a Trillion Minutes of Wearable Data, It Wins 33 of 35 Health Tasks
Pretrained on 2 billion hours of wearable data from 5 million people, a frozen encoder with just a linear head beats supervised baselines on 34 of 35 health tasks.
№33·07-10·23:48🔒 MemberRead article ↗
Every's 9-Person Team Ran GPT-5.6 Sol for a Month: Here's the Verdict
First-hand impressions from four scenarios: coding, writing, knowledge work, and agents
№32·07-10·21:54Read article ↗
Claude Code's Two Knobs, Explained: One for Capability, One for Effort
A ClaudeDevs deep dive untangles two knobs that both seem to promise a better answer: switching models swaps in a different frozen set of weights, while dialing up effort changes how willing it is to read more files, run more tests, and double-check before handing in the result.
№31·07-09·12:14🔒 MemberRead article ↗
Gemma 4 Technical Report: How a Small Open Model Takes On Large Ones with Reasoning and Memory Efficiency
Not a feature list — three engineering tracks turned at once: reasoning lifts intelligence, an efficiency stack cuts cost, native multimodality expands input.
№30·07-09·19:17🔒 MemberRead article ↗
LangChain Tunes the Harness, Not the Model — Nemotron 3 Ultra Closes In on Opus 4.8 at 1/10 the Cost
Three levers — system prompt, tool descriptions, middleware — push the Deep Agents suite from a typical ~0.80 to 0.84, topping out at 0.86 against Opus's 0.87.
№29·07-09·20:40🔒 MemberRead article ↗
AI Self-Improvement Starts Outside the Model: Lilian Weng on Harness Engineering
Former OpenAI safety lead surveys nearly 30 papers: from prompt tweaks to self-modifying code, DGM pushed coding ability from 20% to 50%.
№28·07-09·23:26🔒 MemberRead article ↗
HKUDS Open-Sources AI Tutor DeepTutor — 20K GitHub Stars in 111 Days
Paper shows a 10.8% gain in personalization and 29.4% in reasoning, fully open-source under Apache 2.0
№27·07-08·22:50Read article ↗
Anthropic's Playbook: Fable 5 Advises, Sonnet 5 Foots the Bill
In advisor mode Fable 5 just gives advice; in orchestrator mode it delegates tasks — either way, the cheaper Sonnet 5 ends up doing most of the work
№26·07-08·20:03🔒 MemberRead article ↗
Liquid AI's Antidoom Fixes Reasoning Models' Doom Loops With a Single Token — Now Open-Sourced
By fine-tuning only the single token where the doom loop begins, both models' loop rates drop to around 1%
№25·07-08·21:46🔒 MemberRead article ↗
Cloudflare's New Gateway Lets AI Pay You Automatically to Crawl Your Site
Expanding from charging only AI crawlers to charging any caller, now in early access waitlist
№24·07-07·17:54Read article ↗
Anthropic Discovers a Brain-Like 'Inner Workspace' in Claude — It Evolved, Wasn't Designed
It makes up less than 10% of the model — remove it and Claude can still talk, but its reasoning collapses to zero. Anthropic is already using it to catch fabricated data and spot when Claude senses it's being tested.
№23·07-07·27:57Read article ↗
Dartmouth Put AI Grading to the Test: Students Called It Rigid — But Users Scored Higher
In a 151-student trial, short-answer questions moved scores more than multiple choice, while almost no one touched the AI help sidebar
№22·07-06·24:56Read article ↗
An AI Engineer Built His Nonverbal Autistic Son a Communication App — Speech Jumped 5x, and It Became a Real Business
No funding, no team — built purely to solve his own kid's problem. Clinics and schools started asking to use it anyway.
№21·07-06·20:56Read article ↗
AIEWF closing debate pits agentic coding-loop hype against engineering discipline
A live audience vote couldn't be tallied because the venue lights were too bright to count hands, but a companion survey found 95% of teams already use agents while 59% worry about mounting technical debt.
№20·07-06·21:52Read article ↗
Wealthy American Families Ditch Public School for $75,000-a-Year AI Private Schools
While traditional schools are still figuring out AI, Silicon Valley and Wall Street families are already voting with their wallets.
№19·07-06·21:50Read article ↗
AI Is Torching Junior Dev Jobs: Coding Is Becoming a Basic Skill, Not a Career
US developer employment among 22-to-25-year-olds has fallen 19% in three years, even as new GitHub sign-ups hit their fastest growth ever.
№18·07-05·20:51Read article ↗
NVIDIA Research Unveils HORIZON: Unattended Agents Push Full RTL Chip Design Benchmark to 100% Pass Rate
The paper is the first to run an agent through an entire RTL benchmark suite fully unattended—most tasks clear in two or three rounds, but the hardest one takes 82 iterations
№17·07-05·21:21🔒 MemberRead article ↗
Anthropic Analyzed 400K Claude Code Sessions: Expertise Beats Coding Skill
A 7-month analysis of sessions from 235,000 users: verified experts succeed at nearly double the rate of novices — yet the top 10 professions differ by no more than 7 percentage points.
№16·07-05·22:24🔒 MemberRead article ↗
Dan Koe's Guide to Persuasion: 3 Tensions, 5 Levers — Steal It for Your Copy
Create psychological tension first, then offer a first step too small to refuse — it works for writing, selling, and job hunting alike.
№15·07-05·25:17Read article ↗
Sakana AI Co-Founder: The Ability to Swap Models on Demand Is Itself a Deterrent
On June 12, U.S. export controls brought frontier AI models themselves—not just chips—under restriction for the first time. His answer: master multi-model orchestration.
№14·07-04·21:53🔒 MemberRead article ↗
Working With Claude Fable 5: The Real Skill Is Finding Your Own Unknowns
Anthropic's Thariq argues the quality of your work with Claude Fable 5 hinges on how clearly you can name your own unknowns. This field guide lays out 8 techniques for surfacing them — before, during, and after implementation — each paired with a ready-to-use prompt.
№13·07-04·21:30🔒 MemberRead article ↗
As AI Gets Cheaper, Renting Generic Intelligence Gets Riskier
Investor Chamath Palihapitiya: intelligence is getting cheap like phones, and expert judgment is now available to everyone — the real moat is encoding your proprietary experience into your own system, not renting the same generic AI as your competitors
№12·07-03·23:29🔒 MemberRead article ↗
Every's Head of Consulting Can't Code — She Still Builds Her Own Work Systems With Codex
She used the same pattern to build an email triage tool and a caregiving app for her dad — the method is repeatable, though the full prompts remain unpublished.
№11·07-03·20:05🔒 MemberRead article ↗
Same AI Model, Opposite Outcomes: Why Some Companies Compound Gains While Others Get Nothing
Four implementation principles, backed by real-world data from L'Oréal, Lyft, and Rakuten
№10·07-03·25:41Read article ↗
Alibaba Open-Sources Page Agent: Agents Embedded in the Webpage, Reading Text Instead of Screenshots
MIT-licensed and model-agnostic — works with any OpenAI-compatible text model, though for now it only handles a single page view
№09·07-03·17:47Read article ↗
未来的网站会为每个访客现场拼装:Adobe 演示按意图实时生成专属页面
现场demo:搜'露营'后,咖啡机网站文案产品全变户外主题;技术能落地,客户网站还没规模上线。
№08·07-03·9:50🔒 MemberRead article ↗
The AI Leverage Ladder for PMs: Two Paths to Higher Output, Plus 3 Copy-Paste Prompts
An instructor who has trained 30,000+ PMs breaks down two AI leverage ladders — from copy-paste to end-to-end delivery, and from web prototypes to production PRs.
№07·07-03·13:26🔒 MemberRead article ↗
Claude Code's Official Playbook: 4 Levels of Agent Loops to Unattended
From manual confirmation to fully unattended, the Claude Code team lays out a 4-level loop taxonomy with practical guidance
№06·07-02·21:26🔒 MemberRead article ↗
Meituan Launches LongCat-2.0: 1.6-Trillion-Parameter Model Trained Entirely on Domestic Chips, No Nvidia GPUs
Trained on 50,000+ domestic AI chips and 35 trillion tokens; most benchmarks come from Meituan's own evaluation framework, and the weights aren't truly open for download yet
№05·07-01·22:45Read article ↗
Anthropic Launches Claude Science: An AI Workbench for Scientists with 60+ Built-In Research Skills
Now in open beta. A coordinator agent marshals a team of expert agents to do the work, with a reviewer agent at the end dedicated to catching errors in citations and numbers — compute gets outsourced to AI, but raw data never leaves your local machine.
№04·07-01·22:19Read article ↗
Bridgewater Built a Financial-Filtering Model with 84.7% Accuracy — and Open-Sourced the Method
Partnering with Thinking Machines, they fine-tuned an open-source model on expert-labeled data: 29.8% lower error rate than the best frontier model, at just 1/14 the inference cost
№03·07-01·12:32🔒 MemberRead article ↗
AI Replaced 700 Agents at Klarna. Then Quality Dropped.
Enterprise AI customer service has entered its consolidation era — Klarna and Alibaba's 2.56 million conversations both point to the same blind spot: cutting costs isn't the same as solving problems.
№02·06-30·22:45🔒 MemberRead article ↗
What Is Loop Engineering, Actually?
An Anthropic engineer's methodology for "loop engineering": instead of prompting AI one line at a time, design a self-running loop system
№01·06-29·24:49🔒 MemberRead article ↗
0:000:00