Library / All

All explainers

151 explainers · newest first

Free
Deep Dive
№ 1102
DEEP DIVE
Source · Invest Like The Best⚑ Expert view

Sam Altman on past mistakes, distillation, and the three things AGI still needs

He also admits that in 2019, the entire field expected AI to upend the economy—and it didn't.

Sam AltmanOpenAI07-30 · 12 min read
Free
Research
№ 1101
RESEARCH
Source · BleepingComputer⚑ Research

LayerX finds a new visual fraud attack that tricks AI browser assistants

A custom font and a few lines of CSS are enough—no JavaScript, and no browser vulnerability required.

AI securitybrowser assistants07-30 · 10 min read
Free
Launch
№ 1100
LAUNCH
Source · Replit 官方博客⚑ Product

Replit launches Replit Design: one click for another layout, and a mockup that becomes a working site

Canvas and click-to-edit are free; turning a design into an App starts at Core for $25 a month.

ReplitAI design tools07-30 · 9 min read
Research
🎧№ 1099
RESEARCH
Source · OpenAI⚑ Tutorial

Two API settings, not a new model: OpenAI reports GPT-5.6 Sol jumping from 13.3% to 38.3% on ARC-AGI-3

The model stayed the same; the harness did not—retained private reasoning plus compaction instead of deletion cut output tokens per game to about one-sixth.

ARC-AGI-3harness▶ Play07-30 · 9 min read
Tools
№ 1098
TOOLS
Source · Magnific Academy⚑ Tutorial

Magnific's 38-page handbook for image and video prompts that no longer work on frontier models

Copy-ready prompts, banned-word lists, and camera rules are all there—but the model rankings come from Magnific's own team, with no disclosed test method.

promptingimage generation07-29 · 18 min read
Free
Research
№ 1097
RESEARCH
Source · Anthropic⚑ Research

Claude Mythos finds pure mathematical flaws in encryption schemes, matching elite cryptographers

Neither result threatens anything in production: HAWK is not deployed yet, and the AES work hit only a seven-round reduced version, not the full ten-round standard.

CryptographyAI research07-29 · 13 min read
Tools
🎧№ 1096
TOOLS
Source · GitHub Blog⚑ Tutorial

How GitHub’s official harness guide becomes a clear, repeatable 8-step workflow

He even installs two skills in the piece itself — and that “mostly” at the end of the original title is doing real work.

GitHub CopilotAI coding▶ Play07-29 · 12 min read
Free
Launch
№ 1095
LAUNCH
Source · MCP 官方博客⚑ Product update

Anthropic releases new MCP (Model Context Protocol) spec that lets servers run in the cloud

It's the biggest change since MCP launched, but July 28 isn't a hard cutover — older implementations keep working as before.

MCPProtocol Spec07-29 · 12 min read
Free
Launch
🎧№ 1094
LAUNCH
Source · OpenAI 开发者文档⚑ Product

OpenAI's gpt-transcribe and gpt-live-transcribe halve real-world word error rate, cut price 25%

The real story is in the control test: three older models fed the same background context showed no gains — and two actually got worse.

OpenAISpeech-to-text▶ Play07-29 · 9 min read
Tools
№ 1093
TOOLS
Source · Florian Darroman 播客⚑ Tutorial

How to get AI to recommend your SaaS product and drive real conversions: A GEO playbook from Ranking on AI's founder

Real GEO success is measured in signups and revenue rather than mention screenshots, and it starts by ditching 'what is X' articles to capture high-intent 'best X' queries.

AI visibilitySEO07-28 · 15 min read
Free
Research
🎧№ 1092
RESEARCH
Source · OpenAI Economic Research⚑ Research

OpenAI analyzed 800,000 work-related messages: most of what people use AI for is not what their job was supposed to be

The exact same dataset produces task crossover rates of 43.5% and 65%–82%, differing only in how the denominator is sliced.

OpenAILabor economics▶ Play07-28 · 12 min read
Free
Research
🎧№ 1089
RESEARCH
Source · Kimi K3 技术报告⚑ Research

Kimi K3 technical report: How three architectural changes boosted compute efficiency 2.5x

Published 11 days after launch, the 47-page paper focuses on efficiency, delivering 2.5 times the performance of K2 on the same compute budget.

Kimi K3Moonshot AI▶ Play07-28 · 16 min read
Free
Deep Dive
🎧№ 1088
DEEP DIVE
Source · 陶哲轩个人网站⚑ Expert view

Terence Tao at the International Congress of Mathematicians: AI will generate more proofs, but mathematics won't speed up

Mathematical research is a five-step pipeline, and Terence Tao argues AI only speeds up step one—the remaining four steps grow progressively slower and rely even more heavily on human judgment.

Terence TaoAI and Mathematics▶ Play07-27 · 11 min read
Deep Dive
🎧№ 1087
DEEP DIVE
Source · Anthropic 官方博客⚑ Tutorial

The story behind Claude Design and 10 practical tips from its creator

The three most counterintuitive takeaways: ask for ten variations at once, stick to wireframes when details do not matter, and handle the final mile yourself.

Claude DesignAI design tools▶ Play07-27 · 9 min read
Deep Dive
№ 1086
DEEP DIVE
Source · Dex Horthy(HumanLayer)⚑ Expert view

Harness Engineering isn't enough: Why the software factory model fails

Telemetry tracking 22,000 developers over two years shows a 66% increase in output alongside a 242.7% surge in production incidents per PR.

AI codingSoftware engineering07-27 · 12 min read
Free
Research
№ 1085
RESEARCH
Source · Google Blog / AI & Economy ATLAS v1.0⚑ Research

Google combed through 14.65 million Gemini chats — 86% weren't about work

The median occupation has AI touching just one-fifth of its tasks, 29% of jobs show zero AI use at all, and even in cognitive work, AI carries a task start to finish only 6.5% of the time.

GoogleAI and the Economy07-27 · 14 min read
Deep Dive
🎧№ 1084
DEEP DIVE
Source · Anthropic 官方文档⚑ Tutorial

Anthropic's official guide to prompting Opus 5: subtract, don't add

Eight prompts you can copy straight into your workflow — plus one counterintuitive lesson: telling a review prompt to flag only high-severity issues can genuinely make it report less.

AnthropicClaude Opus 5▶ Play07-25 · 10 min read
Tools
№ 1083
TOOLS
Source · Claude 官方博客⚑ Tutorial

Anthropic turns manual code review into a self-checking, self-fixing skill

Internally, four skills run in relay — catching bugs, cleaning up the diff, booting the app to see it work, and re-checking the design spec whenever a change touches the UI.

Claude CodeSkills07-25 · 9 min read
Free
Launch
№ 1080
LAUNCH
Source · Anthropic⚑ Product

Claude Opus 5 closes the gap with Fable 5 at half the price, but still trails on 5 of 12 benchmarks

The same price now buys roughly double the score, and this release's chart plots cost on the x-axis instead of accuracy — five effort tiers let you dial in exactly how good you want the model and how much to pay for it.

Claude Opus 5Anthropic07-25 · 12 min read
Tools
🎧№ 1078
TOOLS
Source · Thariq / Anthropic⚑ Tutorial

Anthropic slashes Claude Code's system prompt by 80%: six new rules for writing context

Rules gave way to judgment calls and examples gave way to interfaces — one tool description shrank from roughly 9,100 characters to a single sentence and an enum.

Claude CodeContext Engineering▶ Play07-25 · 10 min read
Free
Launch
№ 1077
LAUNCH
Source · Black Forest Labs⚑ Product

Black Forest Labs unveils FLUX 3, a single model for image, video, audio, and robotics

Video access opens for application now; image generation is still weeks away, with open weights coming last. Pricing remains unannounced, and Black Forest Labs itself labels the human-evaluation results as preliminary.

FLUX 3Black Forest Labs07-25 · 9 min read
Free
Deep Dive
🎧№ 1076
DEEP DIVE
Source · Peter McCrory(X)⚑ Expert view

Why US unemployment hasn't risen even as AI races ahead, according to Anthropic's economics chief

His take: AI is amplifying workers rather than replacing them—killing tasks is not the same as killing jobs.

AI and jobsLabor economics▶ Play07-23 · 11 min read
Free
Launch
№ 1075
LAUNCH
Source · OpenAI⚑ Product

OpenAI launches Presence: managing enterprise AI Agents after go-live is now a product

Built for customer-facing and internal workflows, Presence opens with voice and chat first—OpenAI says its own phone line already runs 75% without humans, and a Codex improvement loop cut human handoffs another 15 points in 10 days.

OpenAIPresence07-23 · 8 min read
Deep Dive
🎧№ 1074
DEEP DIVE
Source · Y Combinator⚑ Expert view

YC names the startup ideas it's eager to fund right now

From AI tutors for kids to data centers built at sea — and, for the first time, a request from the sitting US Secretary of the Army.

Y CombinatorStartup Ideas▶ Play07-23 · 10 min read
Free
Deep Dive
№ 1073
DEEP DIVE
Source · TechCrunch Equity⚑ Expert view

Anthropic investor Matt Murphy: revenue tops $47 billion, but the model isn't what wins

The investor who led a $500 million Series D says Anthropic's real moat isn't the model — it's the layer that makes it usable.

AnthropicAI investing07-23 · 11 min read
Research
№ 1071
RESEARCH
Source · Databricks⚑ Company PR

Databricks benchmarks AI coding agents on millions of lines of real code

Swapping the harness around the same model can double the cost, and open-source GLM 5.2 matches Opus 4.8 for 30% less per task — on a benchmark built from Databricks' own merged pull requests, so none of it is searchable online.

Coding agentsBenchmarks07-22 · 9 min read
Tools
№ 1070
TOOLS
Source · Anthropic / ClaudeDevs⚑ Tutorial

Anthropic's six-step playbook for code migration: how Bun moved a million lines from Zig to Rust in two weeks

The templates and prompts behind it are now open source — Bun's migration alone burned through 5.9 billion uncached input tokens, about $165,000 at API list price.

Code migrationClaude Code07-22 · 12 min read
Free
Deep Dive
🎧№ 1068
DEEP DIVE
Source · OpenAI⚑ Company PR

OpenAI test model escapes its sandbox, breaches Hugging Face's systems

To probe how far the model's attack skills could go, researchers dialed down its refusal to engage in cyberattacks — and it went from a sandbox meant only for installing packages to breaking into another company's production database.

AI SecurityOpenAI▶ Play07-22 · 9 min read
Deep Dive
🎧№ 1067
DEEP DIVE
Source · Jeff Huber / Chroma⚑ Expert view

Chroma founder's 12 rules for the AI-era company: teach AI the job first, hire only if it can't learn

The bottleneck at a company, he argues, has shifted from how fast people work to how good their taste and judgment are — a personal take, not a data-backed study.

Organizational ChangeAI Agent▶ Play07-22 · 12 min read
Free
Launch
№ 1066
LAUNCH
Source · Google Blog⚑ Product

Google ships three new Gemini models — the long-awaited 3.5 Pro isn't one of them

All three are Flash models: the flagship writes two-thirds less, the cheapest tier costs 60% more instead of less, and the third is government-only.

GeminiGoogle07-22 · 12 min read
Free
Launch
№ 1065
LAUNCH
Source · Qwen⚑ Product update

Tongyi's Qwen-Image-3.0 turns one 3,000-character prompt into nine infographics and a full newspaper page

The prompt limit jumps from 1K to 4.5K tokens, with legible text down to 10px, and Alibaba Cloud's Bailian platform already offers qwen-image-3.0-pro access, free for a limited time.

Image GenerationText Rendering07-22 · 13 min read
Research
№ 1064
RESEARCH
Source · Memories.ai Research⚑ Research

Memories.ai's O-MARC makes audio-visual AI faster, cheaper, and more accurate

The release ships with a companion benchmark that strips out the audio track and re-runs the test, filtering out questions models can already answer by sight alone.

Video understandingMultimodal compression07-21 · 12 min read
Deep Dive
№ 1062
DEEP DIVE
Source · First Principles · Rippling⚑ Expert view

VC Sarah Guo on AI startups: what to do when a giant could crush you overnight

Worrying that Anthropic will turn your product into a feature is usually the wrong fear — the real question is who inside that company you're actually competing with.

StartupsVenture capital07-21 · 11 min read
Free
Launch
№ 1061
LAUNCH
Source · Linear⚑ Product

Linear's Loops turn loop engineering into a single plain-English prompt

Available on Business and Enterprise plans, billed per run at $0.07 to $0.20.

LinearAgent07-21 · 9 min read
Deep Dive
№ 1060
DEEP DIVE
Source · Hugging Face Blog⚑ Company PR

Hugging Face and Cerebras make Gemma 4 voice chat feel human

It takes a different route than Doubao's or GPT-Live's end-to-end full-duplex systems — the pacing can't quite match theirs, but every piece of the pipeline is open source, swappable, and runs on your own hardware.

Real-time voiceOpen-source pipeline07-21 · 12 min read
Free
Launch
№ 1059
LAUNCH
Source · FunAudioLLM⚑ Product

Tongyi's Qwen-Audio-3.0-TTS speaks 16 languages and 20 dialects from one reference clip

Speaker similarity ranks first across all 16 languages, and cloning still works even with noisy reference audio — but this time Alibaba is opening only the API, not the model weights.

Speech synthesisQwen07-21 · 9 min read
Research
🎧№ 1058
RESEARCH
Source · New Scientist⚑ Expert view

Study finds a sweet spot for AI-assisted writing — not too much, not too little

Three separate trials of the same study all landed on the middle ground — and the columnist who tried ChatGPT on a movie synopsis says he'd still rather write it himself.

AI WritingCreativity▶ Play07-20 · 7 min read
Free
Deep Dive
🎧№ 1056
DEEP DIVE
Source · scaling01

Have Chinese AI models caught up to the US frontier? A look at the real gap

Kimi K3's launch set off claims that China has caught up with the US, but a different measuring stick puts the lag at more than triple — and even flips which side is accelerating faster.

Benchmarking MethodsUS-China AI Gap▶ Play07-20 · 12 min read
Free
Deep Dive
🎧№ 1054
DEEP DIVE
Source · Hugging Face Blog

An AI agent breached Hugging Face's internal network, then commercial LLMs mistook the investigators for attackers

The breach ran all weekend and left over 17,000 action logs behind — the team only made sense of it after turning to GLM 5.2, an open-source model running in a self-hosted environment.

Security IncidentAI Agent▶ Play07-20 · 9 min read
Research
🎧№ 1053
RESEARCH
Source · arXiv⚑ Research

The model you picked through a router might be swapped — one random number proves it

The fingerprint distance between two samples of the same model has a median of 0.140. One API marketed as a proprietary in-house flagship scores 0.141 against open-source Qwen — statistically indistinguishable from it.

Model fingerprintingAPI auditing▶ Play07-20 · 11 min read
Deep Dive
🎧№ 1051
DEEP DIVE
Source · lopopolo/harness-engineering⚑ Tutorial

Google engineer turns a year of harness engineering into 12 rules and a copy-paste playbook

The counterintuitive core: lock the model and agent in place, don't touch a single parameter, and change only the context and tools around them.

harness engineeringAgent engineering▶ Play07-19 · 20 min read
Free
Launch
🎧№ 1050
LAUNCH
Source · 商汤 SenseNova⚑ Product

SenseTime launches SenseNova U1 Pro: native 8K images that stay sharp even when you zoom into dense small text

At WAIC, the model showed off a 4:1 ink-wash scroll and a 22-panel storyboard sequence. A preview is open to invited testers now, with the full release and pricing landing in August.

SenseTimeMultimodal▶ Play07-19 · 9 min read
Deep Dive
🎧№ 1049
DEEP DIVE
Source · Anthropic⚑ Expert view

Inside Anthropic: the harness wrapped around agents is coming apart

Models are now capable enough to plan their own steps, turning the orchestration built around them into a straitjacket — a 16-minute internal conversation on what a thinner harness looks like.

▶ Play07-18 · 5 min read
Deep Dive
🎧№ 1048
DEEP DIVE
Source · Cerebras⚑ Tutorial

Cerebras details how it built a knowledge base that leaves data where it lives and lets AI fetch it

Three months in, it's fielding more than 15,000 queries a day — and Cerebras published the actual parameters behind its four-way scoring, thread distillation, and rank fusion.

Knowledge baseRetrieval▶ Play07-18 · 12 min read
Free
Deep Dive
🎧№ 1047
DEEP DIVE
Source · a16z⚑ Expert view

a16z: the software selloff is only killing companies without a moat

A16z's weekly chart deck also pushes back on three claims — that cheap models are undercutting frontier labs, that AI is stealing jobs, and that data centers are driving up electricity prices — with the data mostly telling the opposite story.

AI economySoftware stocks▶ Play07-18 · 8 min read
Deep Dive
🎧№ 1046
DEEP DIVE
Source · Tim Ferriss 博客⚑ Expert view

AI can copy your work in seconds — but not these eight things

Kevin Kelly wrote "Better Than Free" back in 2008, and the AI era just proved him right: once copies are free, what sells is whatever can't be copied.

Creator EconomyKevin Kelly▶ Play07-18 · 7 min read
Tools
🎧№ 1045
TOOLS
Source · DEV · Google AI⚑ Tutorial

Google's gemma-trainer skill fine-tunes a custom Gemma right on your own machine

This open-source training playbook locks down all three fine-tuning paths — supervised fine-tuning, preference alignment, and reward scoring — plus the LoRA parameters, pairs with Unsloth, and runs on a consumer GPU with just 8GB of VRAM.

GemmaLocal fine-tuning▶ Play07-18 · 8 min read
Deep Dive
🎧№ 1044
DEEP DIVE
Source · Raft

The CTO is an AI too: how a nearly all-AI team shipped a product feature

Of the team's 15 members, only the founder is human — the rest are AI. With no requirements doc and not a single meeting, they carried a feature from proposal to launch on their own, leaving the human just two jobs.

AI AgentAI Collaboration▶ Play07-17 · 8 min read
Free
Launch
№ 1043
LAUNCH
Source · Bolt.new⚑ Product

Bolt Slides turns AI-generated decks into live web pages — 3D flythroughs, real-time whiteboards, filterable charts, one link to demo

AI-generated slides used to export as flat, dead images. Bolt Slides turns every page back into a working website — 3D scenes you can spin, calculators you can click, whiteboards that take live votes.

Bolt SlidesOpen Source07-17 · 9 min read
Research
🎧№ 1042
RESEARCH
Source · Impossible Research⚑ Research

A single wrapper pushes Opus 4.8 and Fable 5 to 99% on ARC-AGI-3

A harness called Schema has models turn each game's rules into a runnable, verified program before making a move. Across 25 public rounds it self-reported 98.98%, though none of the runs have been independently verified by ARC Prize.

ARC-AGI-3World Model▶ Play07-17 · 12 min read
Free
Launch
🎧№ 1041
LAUNCH
Source · LM Studio⚑ Product

LM Studio launches Bionic, a Codex-style agent for open models

Run it locally for free with your data staying on-device, or switch to the cloud for more horsepower — zero data retention by default. The preview is live today on Mac and Windows.

LM StudioLocal AI▶ Play07-17 · 6 min read
Free
Launch
№ 1040
LAUNCH
Source · LiveKit⚑ Product

LiveKit optimizes Gemma 4 31B for real-time voice, 5x faster and 6x cheaper than GPT-4.1

Same answer quality, 5.2x lower turn-taking latency, and 6x lower cost than GPT-4.1 — pointing a voice agent at it takes a one-line code change. Figures are from LiveKit's own benchmarks.

LiveKitGemma 407-17 · 6 min read
Free
Launch
№ 1038
LAUNCH
Source · Decart⚑ Product

Decart ships Lucy 2.5 for live video edits with sub-40ms latency

Add objects, swap backgrounds, and apply effects mid-stream—Decart claims continuous generation under 40ms at 100FPS, with a try-it link at lucy.decart.ai.

DecartLucy 2.507-17 · 6 min read
Free
Launch
🎧№ 1037
LAUNCH
Source · Kimi 官方博客⚑ Product

Moonshot AI ships Kimi K3, the world's first 3-trillion-parameter open model

2.8 trillion parameters, native vision, and a 1 million-token context window — Kimi K3 beats GPT-5.6 Sol on several agent benchmarks, with the app and API live today

Kimi K3Moonshot AI▶ Play07-17 · 9 min read
Free
Business
№ 1036
BUSINESS
Source · 404 Media

Suno training library exposed in hack: source code shows ~380,000 hours scraped from YouTube Music and more

A worm-led breach of Suno spilled source code that listed which sites were scraped for training audio and for how many hours, along with user data the company says it does not need to notify users about one by one.

Sunotraining data07-17 · 8 min read
Free
Research
№ 1035
RESEARCH
Source · Design Arena

GPT-5.6 Sol tops front-end design leaderboard by dodging AI's usual clichés

Design Arena combed through a thousand of the model's web pages and found a gap in the design space — right where clichés like purple gradients usually cluster.

GPT-5.6 SolAI web design07-16 · 6 min read
Deep Dive
🎧№ 1034
DEEP DIVE
Source · Google Search Central⚑ Tutorial

How Should Websites Optimize in the AI Search Era? Google's Official SEO Guide

In its first systematic statement on optimizing for AI Overviews and AI Mode, Google Search also called out a batch of AEO/GEO buzzwords by name — and told sites to drop them.

Google SearchSEO▶ Play07-16 · 8 min read
Free
Launch
№ 1032
LAUNCH
Source · Pixel Point⚑ Product

Web micro-animations hitch on every interaction? AVAL wants prerendered video to understand hover, click, and success

Pain first: small 3D icons need to move and react—ordinary video often stutters. AVAL is a format and player for that short motion.

AVALinteractive video07-16 · 6 min read
Free
Launch
№ 1031
LAUNCH
Source · xAI GitHub⚑ Expert view

Inside Grok Build’s Source: After Reports of Full-Repo Uploads, Musk Open-Sourced Grok Build

First a privacy failure, then delete data, turn off defaults, and open-source the whole tree. Below we open the source: size, prompts, tools, editing, uploads, memory, and safety.

Grok BuildxAI07-16 · 10 min read
Free
Launch
№ 1029
LAUNCH
Source · Thinking Machines Lab⚑ Product

Mira Murati’s First Open Model Activates Just 41 Billion Parameters per token

Thinking Machines Lab, founded by former OpenAI CTO Mira Murati, has released Inkling, its first large model trained in-house. It can read text, interpret images, listen to audio, write code, and use tools. Its full weights are available to download, and developers can continue training it on Tinker.

InklingThinking Machines Lab07-16 · 12 min read
Deep Dive
🎧№ 1021
DEEP DIVE
Source · George Sivulka / X⚑ Expert view

You just hired a million bad employees

Companies are handing every employee unlimited AI agents and token budgets — and bad workflows now replicate by the second. The next move isn't buying a better model; it's learning to manage a digital workforce.

AI AgentToken▶ Play07-16 · 11 min read
Free
Research
№ 1027
RESEARCH
Source · MIT News / Science⚑ Research

MIT Builds a Robotic Bird That Swims and Flies Without Paddling Feet

A 250-gram robot crosses water and air with the same flexible wings, taking off from a lake at a 70° pitch after 8–10 wingbeats

MITBio-inspired robotics07-15 · 8 min read
Free
Launch
№ 1026
LAUNCH
Source · Google Developers Blog⚑ Product

Google runs Gemma 4 on Pixel 10: chat, vision, and settings changes at 30,000 feet offline

At Google I/O India, Tensor and Pixel jointly showed lightweight Gemma 4 on the phone TPU—on-device, no cloud. Developers can apply for Tensor SDK.

Google TensorPixel 1007-15 · 7 min read
Deep Dive
🎧№ 1024
DEEP DIVE
Source · Databricks Blog⚑ Company PR

How much did this order really make? Retail finance turns to AI

Customers buy online, pick up in store, and return across locations—one purchase, many ledgers. Genie is framed as a finance sidekick for true profit, trapped cash, and fewer markdowns.

DatabricksGenie▶ Play07-15 · 10 min
Tools
№ 1023
TOOLS
Source · pols.dev⚑ Tutorial

An 80k-Character Anti-Slop Design Law: Drop Slop.md Into Your Agent

pols.dev ships a ~87KB Markdown rulebook naming AI UI tells and positive recipes. Install by agent path, or download the file—not a website generator.

Slop.mdpols.dev07-15 · 10 min
Tools
№ 1022
TOOLS
Source · Impeccable⚑ Tutorial

Design Skills That Strip AI Taste From Frontend: Impeccable

One core skill, 23 design commands, an anti-pattern list, plus CI-ready slop detection. The main site shows the workflow; /slop lays out the "looks AI-made" UI tells.

ImpeccableAI design07-15 · 10 min
Free
Research
🎧№ 1020
RESEARCH
Source · PrismML⚑ Research

PrismML Stuffs a 27B Model into Your iPhone — With Barely Any IQ Drop

It squeezes a ~54GB 27B model down to about 3.9–5.9GB: it runs on-device, and average scores still hold about nine-tenths. Core value first; technical detail and the fine print come later.

Bonsai 27BOn-device AI▶ Play07-15 · 12 min
Free
Deep Dive
🎧№ 1019
DEEP DIVE
Source · Demis Hassabis · X⚑ Expert view

DeepMind's Demis Hassabis: give frontier models a 30-day safety check before release — or keep them off the market

Hassabis wants the US to stand up a FINRA-style body that defines frontier models with a moving benchmark — a voluntary protocol now, a hard gate to market later.

Demis HassabisDeepMind▶ Play07-14 · 6 min read
Free
Deep Dive
🎧№ 1018
DEEP DIVE
Source · AI as Normal Technology⚑ Expert view

Princeton professor at ICML on AI and work: how should individuals adapt?

In ~24 months, model capability rose sharply; SAGE’s composite reliability metric rose only 5–10 percentage points

AI as Normal TechnologyAI reliability gap▶ Play07-14 · 12 min read
Deep Dive
🎧№ 1017
DEEP DIVE
Source · Anthropic⚑ Company PR

Anthropic's playbook for AI-native startups: four stages, graduation criteria, and ready-to-copy prompts

Anthropic's 36-page playbook breaks down the graduation bar and common pitfalls for four startup stages, complete with matching Claude prompts you can copy straight in.

Startup methodologyClaude Code▶ Play07-14 · 14 min read
Deep Dive
🎧№ 1016
DEEP DIVE
Source · DoorDash Engineering Blog⚑ Company PR

DoorDash's AI Shopping Assistant: It's Not the Model, It's the Tools and Memory

After the memory system launched, grocery checkout conversion rose about 24%, and automated evals scaled daily test volume from 1 human-reviewed case to 2,000+.

AI Shopping AssistantAgent Architecture▶ Play07-14 · 14 min read
Deep Dive
🎧№ 1015
DEEP DIVE
Source · ByteByteGo

How Microsoft Ships AI Agents at Enterprise Scale: Breaking a Production Agent into 5 Layers

Retrieval becomes a sub-agent that plans and retries; evaluation upgrades from "does it run" to "did it do it right"

AI Agent ArchitectureMicrosoft Foundry▶ Play07-14 · 12 min read
Free
Research
🎧№ 1014
RESEARCH
Source · Anthropic⚑ Research

Anthropic Analyzed 300,000 Conversations: Claude's Values Shift by Language

Across 3 models and 20 languages: English is the most cautious and in-depth, Russian the most exacting, Hindi the warmest, and Chinese sits closest to the global average

Anthropic ResearchClaude Values▶ Play07-14 · 9 min read
Free
Research
🎧№ 1013
RESEARCH
Source · Sakana AI / Nature Communications⚑ Research

Sakana AI's Brainless 3D Cell Bricks Self-Organize With Biology-Like Swarm Intelligence

With no central brain in charge, nearly 200 simple smart cubes figure out what shape they've formed just by talking to their neighbors—and can even sense where to "regrow" after damage. The self-recognition part already works on physical bricks; damage localization and regeneration still happen mostly in simulation.

Sakana AISwarm Intelligence▶ Play07-14 · 9 min read
Free
Deep Dive
🎧№ 1012
DEEP DIVE
Source · Satya Nadella on X⚑ Expert view

Microsoft CEO Satya Nadella: Companies Face a ‘Reverse Information Paradox,’ Paying for AI While Handing Over Their Internal Knowledge and Experience

Companies pay for intelligence twice: once in model fees, and again in the proprietary know-how required to make the model genuinely useful.

Satya NadellaEnterprise AI▶ Play07-13 · 7 min read
Research
🎧№ 1010
RESEARCH
Source · arXiv · Meta AI⚑ Research

Meta AI's Proactive Memory Agent Teaches Models When Not to Remind You, Boosting Terminal-Bench Accuracy by 8.3 Points

The paper claims the code is open-sourced — but the repo turns out to be empty, without a single commit ever pushed.

AI AgentLong-Task Memory▶ Play07-13 · 9 min read
Deep Dive
🎧№ 1009
DEEP DIVE
Source · Ploy Blog⚑ Company PR

Ploy Switched Production AI from Opus 4.8 to GPT-5.6 Sol: Cutting Latency by Over Half and Costs by 27% Without Losing Quality

Swapping models isn't just swapping an API: eval frameworks, tool parameters, caching, and reasoning traces — four invisible pitfalls, unpacked and fixed one by one

GPT-5.6 SolAI Agent Engineering▶ Play07-13 · 9 min read
Research
🎧№ 1008
RESEARCH
Source · Anthropic 官方博客⚑ Research

Claude Cowork's Top Use Case Is Office Admin (33.4%), Not Coding (8.7%): Interface Shapes How AI Gets Used

Based on 1.2M+ conversations across 600,000+ organizations: content creation ranks second at 16.4%, together accounting for nearly half of all usage.

Claude CoworkAnthropic▶ Play07-13 · 7 min read
Deep Dive
🎧№ 1007
DEEP DIVE
Source · Varick Agents⚑ Expert view

Low-Margin Industries Are AI's Biggest Winners

Manufacturing, logistics, warehousing, and labor services have been stuck at single-digit margins for years — cutting coordination costs alone can multiply their profits.

AI TransformationLow-Margin Industries▶ Play07-12 · 6 min read
Free
Research
🎧№ 1006
RESEARCH
Source · VentureBeat⚑ Research

The Package AI Told You to Install Doesn't Exist — Hackers Already Named Their Malware After It

576K samples, 16 models tested: 19.7% of AI-recommended packages are hallucinations, and 43% keep generating the same fake name.

AI SecuritySoftware Supply Chain▶ Play07-12 · 9 min read
Deep Dive
🎧№ 1005
DEEP DIVE
Source · Reddit r/codex · OpenAI 官方 AMA⚑ Community

OpenAI Staff Reveal How They Really Use GPT-5.6: Unlimited Quota, Not Always Maxed Out

The team dodged questions about benchmark gaming, faced backlash from longtime users over the desktop app merger, and admitted they're "still figuring it out."

OpenAI CodexGPT-5.6▶ Play07-11 · 9 min read
Free
Deep Dive
🎧№ 1004
DEEP DIVE
Source · Thinking Machines Lab⚑ Company PR

Mira Murati, OpenAI's Former CTO: AI Should Amplify Humans, Not Replace Them

From tacit knowledge to interaction bandwidth to model alignment — why AI's progress still can't do without humans.

Thinking Machines LabAI Alignment▶ Play07-11 · 6 min read
Deep Dive
🎧№ 1003
DEEP DIVE
Source · Sierra⚑ Company PR

One AI Per Department Backfired — Sierra's Single Assistant Now Handles 70% of Its Code

A retrospective: since launch, Pinecone has run 75,000 sessions, served 600+ employees, and connected 37 internal systems via MCP Gateway.

Enterprise AI AgentSierra▶ Play07-11 · 8 min read
Deep Dive
🎧№ 1002
DEEP DIVE
Source · ChatGPT Learn⚑ Tutorial

OpenAI's Official Guide: 9 Copy-Ready Codex Prompt Workflows for Fixing Bugs, Turning Screenshots into Prototypes, and Cloud Refactors

OpenAI consolidates prompting tips scattered across its product pages into one framework — goal, context, output, constraints — plus dedicated Codex workflow examples.

Prompt EngineeringChatGPT▶ Play07-11 · 9 min read
Research
🎧№ 1001
RESEARCH
Source · EPFL 项目主页⚑ Research

What Does Each Brain Region Like to Watch? EPFL Evolved AI-Generated Videos to Find Out

All results are computer simulation predictions from a brain "digital twin" model, not yet validated with real human brain imaging.

NeuroscienceEvolutionary Algorithm▶ Play07-11 · 9 min read
Free
Launch
🎧№ 1000
LAUNCH
Source · LangChain Blog⚑ Product

LangChain Ships OpenWiki 0.1.0: Proactive Memory for AI Agents, Auto-Syncing Gmail, Notion, Git, and X into a Local Wiki

No manual context-feeding required — the local Markdown wiki refreshes on a set schedule; a Slack connector is coming soon.

LangChainOpenWiki▶ Play07-11 · 6 min read
Free
Launch
🎧№ 114
LAUNCH
Source · Google Cloud⚑ Product

Google Cloud Puts a Proxy Model Inside AlloyDB: In-Database AI Inference, Up to 23,000x Faster

Local small models replace cloud AI calls — the numbers come from Google's internal testing; proxy models are currently limited to the ai.if function and still in preview

AlloyDBProxy Models▶ Play07-10 · 8 min read
Free
Launch
№ 113
LAUNCH
Source · 1X Tech⚑ Product

1X Gives NEO a New Hand That Can Feel What It's Holding

Every joint in the tendon-driven hand can sense external force, fingertip positioning is accurate to ±0.2mm, and a dedicated production line is planned for an annual capacity of 10,000 units.

Humanoid RobotsDexterous Hand07-10 · 6 min read
Research
🎧№ 112
RESEARCH
Source · Google Research⚑ Research

Google Research Unveils SensorFM: Trained on a Trillion Minutes of Wearable Data, It Wins 33 of 35 Health Tasks

Pretrained on 2 billion hours of wearable data from 5 million people, a frozen encoder with just a linear head beats supervised baselines on 34 of 35 health tasks.

SensorFMWearable Health Data▶ Play07-10 · 8 min read
Free
Deep Dive
🎧№ 111
DEEP DIVE
Source · Every

Every's 9-Person Team Ran GPT-5.6 Sol for a Month: Here's the Verdict

First-hand impressions from four scenarios: coding, writing, knowledge work, and agents

GPT-5.6 SolOpenAI▶ Play07-10 · 9 min read
Free
Launch
№ 110
LAUNCH
Source · OpenAI 官方博客⚑ Product

OpenAI Launches GPT-5.6: Scores Close In on Claude Fable 5 at Half the Cost, Merges ChatGPT and Codex

Sol, Terra, and Luna tiers ship alongside a 4-agent parallel ultra mode; ChatGPT Work launches the same day as Codex folds into ChatGPT. Includes the full bilingual launch keynote.

GPT-5.6OpenAI07-10 · 12 min read
Free
Launch
№ 109
LAUNCH
Source · OpenAI⚑ Product

OpenAI Launches ChatGPT Work, Powered by GPT-5.6, to Auto-Generate Docs, Sheets, and Slides Across Tools

Available today on desktop for all tiers, with web and mobile rolling out to remaining plans in the coming days.

ChatGPT WorkGPT-5.607-10 · 3 min read
Free
Launch
№ 108
LAUNCH
Source · Anthropic⚑ Product

Anthropic Adds Reflect to Claude: Screen-Time-Style Insight Into Your AI Habits

One click in Settings pulls up a report — what you talked about, when you chatted most, what you kept asking it to do, all laid out. Memory has to be turned on first.

AnthropicClaude07-10 · 6 min read
Free
Launch
№ 107
LAUNCH
Source · Cloudflare⚑ Product

Cloudflare Ships Drop: Drag a Folder In, Get a Live Site Worldwide — No Signup Required

A public link in seconds — try it first, log in later to claim it. Unclaimed drops expire in about an hour.

CloudflareDrop07-09 · 7 min read
Deep Dive
🎧№ 106
DEEP DIVE
Source · ClaudeDevs(X)/ Lydia Hallie⚑ Expert view

Claude Code's Two Knobs, Explained: One for Capability, One for Effort

A ClaudeDevs deep dive untangles two knobs that both seem to promise a better answer: switching models swaps in a different frozen set of weights, while dialing up effort changes how willing it is to read more files, run more tests, and double-check before handing in the result.

Claude CodeModel Selection▶ Play07-09 · 9 min read
Free
Launch
№ 105
LAUNCH
Source · Meta AI / Alexandr Wang⚑ Product

Meta Releases Muse Spark 1.1, Rivaling GPT-5.5 and Opus 4.8 on Agentic Benchmarks

Features: orchestrator/sub-agent coordination, million-token context, desktop/browser/mobile control, coding and multimodal; API pricing $1.25 input / $4.25 output per million tokens

MetaMuse Spark07-09 · 10 min read
Research
🎧№ 104
RESEARCH
Source · arXiv / Google DeepMind⚑ Research

Gemma 4 Technical Report: How a Small Open Model Takes On Large Ones with Reasoning and Memory Efficiency

Not a feature list — three engineering tracks turned at once: reasoning lifts intelligence, an efficiency stack cuts cost, native multimodality expands input.

Gemma 4Open-Source LLMs▶ Play07-09 · 12 min read
Research
🎧№ 103
RESEARCH
Source · LangChain Blog⚑ Research

LangChain Tunes the Harness, Not the Model — Nemotron 3 Ultra Closes In on Opus 4.8 at 1/10 the Cost

Three levers — system prompt, tool descriptions, middleware — push the Deep Agents suite from a typical ~0.80 to 0.84, topping out at 0.86 against Opus's 0.87.

LangChainNemotron▶ Play07-09 · 9 min read
Research
🎧№ 102
RESEARCH
Source · Lil'Log⚑ Research

AI Self-Improvement Starts Outside the Model: Lilian Weng on Harness Engineering

Former OpenAI safety lead surveys nearly 30 papers: from prompt tweaks to self-modifying code, DGM pushed coding ability from 20% to 50%.

Harness EngineeringRecursive Self-Improvement▶ Play07-09 · 13 min read
Free
Launch
№ 101
LAUNCH
Source · xAI⚑ Product

xAI Launches Grok 4.5: Rivals Opus 4.8, ~2x Faster Output at a Quarter of the Price

Ranks 3rd on DeepSWE, 4th overall — faster than Opus 4.8 and notably cheaper.

Grok 4.5xAI07-09 · 6 min read
Free
Launch
№ 100
LAUNCH
Source · OpenAI⚑ Product

OpenAI Ships GPT-Live: No More Taking Turns With AI — Hard Questions Auto-Route to GPT-5.5

The first truly full-duplex voice model launches today, and it can hand off complex tasks to GPT-5.5 in real time.

GPT-LiveChatGPT Voice07-09 · 6 min read
Free
Launch
№ 099
LAUNCH
Source · 字节跳动Seed⚑ Product

ByteDance Launches Seedream 5.0 Pro: One Image, a Dozen-Plus Independent Layers — Dense Charts and Real Scenes in a Single Generation

Adds click-and-circle point editing plus native rendering for a dozen-plus languages; now live on Volcano Engine, rolling out to Doubao and Jimeng next

Seedream 5.0 ProByteDance Seed07-08 · 6 min read
Free
Launch
🎧№ 098
LAUNCH
Source · GitHub / arXiv(HKUDS 官方仓库与论文)⚑ Research

HKUDS Open-Sources AI Tutor DeepTutor — 20K GitHub Stars in 111 Days

Paper shows a 10.8% gain in personalization and 29.4% in reasoning, fully open-source under Apache 2.0

DeepTutorHKUDS Open-Source Project▶ Play07-08 · 12 min read
Deep Dive
🎧№ 097
DEEP DIVE
Source · ClaudeDevs(X)⚑ Expert view

Anthropic's Playbook: Fable 5 Advises, Sonnet 5 Foots the Bill

In advisor mode Fable 5 just gives advice; in orchestrator mode it delegates tasks — either way, the cheaper Sonnet 5 ends up doing most of the work

Fable 5Multi-Model Collaboration▶ Play07-08 · 6 min read
Research
🎧№ 096
RESEARCH
Source · Liquid AI⚑ Research

Liquid AI's Antidoom Fixes Reasoning Models' Doom Loops With a Single Token — Now Open-Sourced

By fine-tuning only the single token where the doom loop begins, both models' loop rates drop to around 1%

Liquid AIReasoning Model Doom Loops▶ Play07-08 · 8 min read
Free
Launch
№ 095
LAUNCH
Source · Meta AI Blog⚑ Product

Meta's Muse Image Can Research and Self-Correct — Now Ranks #2 Worldwide

Sibling model Muse Video also debuts with native audio support built in, creator access coming soon.

MetaMuse Image07-08 · 7 min read
Free
Launch
🎧№ 094
LAUNCH
Source · Cloudflare Blog⚑ Product

Cloudflare's New Gateway Lets AI Pay You Automatically to Crawl Your Site

Expanding from charging only AI crawlers to charging any caller, now in early access waitlist

Cloudflarex402 Protocol▶ Play07-07 · 6 min read
Free
Launch
№ 093
LAUNCH
Source · MarkTechPost⚑ Product update

OpenAI Ships GPT-Realtime-2.1 and GPT-Realtime-2.1-mini: Mini Model Gains Reasoning, Cuts Latency 25%

Before calling a tool, the model now says "let me check that" — so calls never go cold waiting

OpenAIRealtime API07-07 · 6 min read
Free
Business
№ 092
BUSINESS
Source · Anthropic 官网⚑ Company PR

Alberta, Canada Used Claude Code for a Security Audit — 466 Million Lines Scanned in 20 Hours, Down From 6.5 Years

50 autonomous agents worked in parallel across 27 provincial departments and 3,400 code repositories — even fixing vulnerabilities and rewriting legacy systems on their own.

Claude CodeCybersecurity07-07 · 5 min read
Free
Launch
№ 091
LAUNCH
Source · Cloudflare Blog⚑ Product update

Cloudflare Workers Cache: Every Worker Gets Built-In Caching, and Hits Cost Zero CPU

One line of Wrangler config plus standard HTTP cache headers — and caching can now sit between any two points inside a Worker.

Cloudflare WorkersEdge Caching07-07 · 7 min read
Free
Deep Dive
№ 090
DEEP DIVE
Source · Anthropic⚑ Company PR

The Untold Story of Claude Code: Boris Cherny's Two-Day Demo Now Writes 100% of His Code

16 insiders recount the journey from wrestling with diffs to a two-week sprint launch — and how engineers stopped writing code by hand

Claude CodeOral History07-07 · 13 min read
Deep Dive
№ 089
DEEP DIVE
Source · AI Engineer 大会 / Anthropic⚑ Tutorial

Claude Fable 5 in Practice: 8 "Find the Unknown" Prompt Patterns, from Blind-Spot Scans to Post-Read Quizzes

Claude Code team's Thariq Shihipar at a conference talk: the bottleneck for new models is no longer the model itself, but whether you can articulate your own unknowns

Claude Fable 5Prompting Techniques07-07 · 11 min read
Free
Research
🎧№ 088
RESEARCH
Source · Anthropic⚑ Research

Anthropic Discovers a Brain-Like 'Inner Workspace' in Claude — It Evolved, Wasn't Designed

It makes up less than 10% of the model — remove it and Claude can still talk, but its reasoning collapses to zero. Anthropic is already using it to catch fabricated data and spot when Claude senses it's being tested.

Interpretability ResearchClaude▶ Play07-07 · 8 min read
Free
Research
🎧№ 087
RESEARCH
Source · iTextbooks'26 论文⚑ Research

Dartmouth Put AI Grading to the Test: Students Called It Rigid — But Users Scored Higher

In a 151-student trial, short-answer questions moved scores more than multiple choice, while almost no one touched the AI help sidebar

AI in EdTechFormative Assessment▶ Play07-06 · 8 min read
Free
Launch
🎧№ 086
LAUNCH
Source · Extelligence(Substack)⚑ Product

An AI Engineer Built His Nonverbal Autistic Son a Communication App — Speech Jumped 5x, and It Became a Real Business

No funding, no team — built purely to solve his own kid's problem. Clinics and schools started asking to use it anyway.

AAC CommunicationAutism Intervention▶ Play07-06 · 7 min read
Free
Deep Dive
🎧№ 085
DEEP DIVE
Source · Latent.Space⚑ Community

AIEWF closing debate pits agentic coding-loop hype against engineering discipline

A live audience vote couldn't be tallied because the venue lights were too bright to count hands, but a companion survey found 95% of teams already use agents while 59% worry about mounting technical debt.

AIEWFagentic loop▶ Play07-06 · 5 min read
Free
Deep Dive
🎧№ 084
DEEP DIVE
Source · The Decoder

Wealthy American Families Ditch Public School for $75,000-a-Year AI Private Schools

While traditional schools are still figuring out AI, Silicon Valley and Wall Street families are already voting with their wallets.

AI EducationEducation Gap▶ Play07-06 · 4 min read
Free
Research
🎧№ 083
RESEARCH
Source · Seldo.com(Laurie Voss)⚑ Research

AI Is Torching Junior Dev Jobs: Coding Is Becoming a Basic Skill, Not a Career

US developer employment among 22-to-25-year-olds has fallen 19% in three years, even as new GitHub sign-ups hit their fastest growth ever.

AI Job DisruptionJunior Developers▶ Play07-05 · 9 min read
Research
🎧№ 082
RESEARCH
Source · arXiv · NVIDIA Research⚑ Research

NVIDIA Research Unveils HORIZON: Unattended Agents Push Full RTL Chip Design Benchmark to 100% Pass Rate

The paper is the first to run an agent through an entire RTL benchmark suite fully unattended—most tasks clear in two or three rounds, but the hardest one takes 82 iterations

NVIDIA ResearchAgentic Workflows▶ Play07-05 · 12 min read
Research
🎧№ 081
RESEARCH
Source · Anthropic⚑ Research

Anthropic Analyzed 400K Claude Code Sessions: Expertise Beats Coding Skill

A 7-month analysis of sessions from 235,000 users: verified experts succeed at nearly double the rate of novices — yet the top 10 professions differ by no more than 7 percentage points.

Claude CodeAnthropic Research▶ Play07-05 · 7 min read
Free
Deep Dive
🎧№ 080
DEEP DIVE
Source · Dan Koe(X 长文)⚑ Expert view

Dan Koe's Guide to Persuasion: 3 Tensions, 5 Levers — Steal It for Your Copy

Create psychological tension first, then offer a first step too small to refuse — it works for writing, selling, and job hunting alike.

Human InsightPersuasion Framework▶ Play07-05 · 8 min read
Deep Dive
🎧№ 079
DEEP DIVE
Source · Yomiuri Shimbun⚑ Expert view

Sakana AI Co-Founder: The Ability to Swap Models on Demand Is Itself a Deterrent

On June 12, U.S. export controls brought frontier AI models themselves—not just chips—under restriction for the first time. His answer: master multi-model orchestration.

AI GeopoliticsExport Controls▶ Play07-04 · 6 min read
Deep Dive
🎧№ 078
DEEP DIVE
Source · X · Thariq⚑ Expert view

Working With Claude Fable 5: The Real Skill Is Finding Your Own Unknowns

Anthropic's Thariq argues the quality of your work with Claude Fable 5 hinges on how clearly you can name your own unknowns. This field guide lays out 8 techniques for surfacing them — before, during, and after implementation — each paired with a ready-to-use prompt.

Claude Fable 5Agentic Coding▶ Play07-04 · 9 min read
Deep Dive
🎧№ 077
DEEP DIVE
Source · Chamath Palihapitiya · X⚑ Expert view

As AI Gets Cheaper, Renting Generic Intelligence Gets Riskier

Investor Chamath Palihapitiya: intelligence is getting cheap like phones, and expert judgment is now available to everyone — the real moat is encoding your proprietary experience into your own system, not renting the same generic AI as your competitors

Falling AI CostsDemocratized Expert Judgment▶ Play07-03 · 5 min read
Free
Business
№ 076
BUSINESS
Source · 微软官方博客⚑ Company PR

Microsoft Launches Frontier Company, Betting $2.5B and 6,000 Experts on Customer AI Transformation

Microsoft pledges customer data won't train models that erode their competitive edge — the platform lets enterprises switch freely between AI models with no vendor lock-in

MicrosoftEnterprise AI Transformation07-03 · 6 min read
Deep Dive
🎧№ 075
DEEP DIVE
Source · Every⚑ Tutorial

Every's Head of Consulting Can't Code — She Still Builds Her Own Work Systems With Codex

She used the same pattern to build an email triage tool and a caregiving app for her dad — the method is repeatable, though the full prompts remain unpublished.

OpenAI CodexLoop Workflow▶ Play07-03 · 9 min read
Free
Business
🎧№ 074
BUSINESS
Source · Anthropic⚑ Company PR

Same AI Model, Opposite Outcomes: Why Some Companies Compound Gains While Others Get Nothing

Four implementation principles, backed by real-world data from L'Oréal, Lyft, and Rakuten

Enterprise AI TransformationClaude Cowork▶ Play07-03 · 13 min read
Deep Dive
№ 073
DEEP DIVE
Source · Anthropic⚑ Tutorial

Anthropic's New AI Agent Architecture Guide: Ask 3 Questions Before Going Multi-Agent

Multi-agent systems can boost performance by 90.2%, but they also cost 10-15x more in tokens. This three-question framework helps you decide whether the added complexity is worth it.

AI Agent ArchitectureMulti-Agent Systems07-03 · 9 min read
Free
Launch
🎧№ 072
LAUNCH
Source · MarkTechPost⚑ Product

Alibaba Open-Sources Page Agent: Agents Embedded in the Webpage, Reading Text Instead of Screenshots

MIT-licensed and model-agnostic — works with any OpenAI-compatible text model, though for now it only handles a single page view

Page AgentAlibaba▶ Play07-03 · 5 min read
Deep Dive
🎧№ 071
DEEP DIVE
Source · Latent.Space

未来的网站会为每个访客现场拼装:Adobe 演示按意图实时生成专属页面

现场demo:搜'露营'后,咖啡机网站文案产品全变户外主题;技术能落地,客户网站还没规模上线。

Adobe智能体网站▶ Play07-03 · 5 min read
Deep Dive
🎧№ 070
DEEP DIVE
Source · Lenny's Newsletter⚑ Tutorial

The AI Leverage Ladder for PMs: Two Paths to Higher Output, Plus 3 Copy-Paste Prompts

An instructor who has trained 30,000+ PMs breaks down two AI leverage ladders — from copy-paste to end-to-end delivery, and from web prototypes to production PRs.

Product ManagerAI Workflow▶ Play07-03 · 9 min read
Free
Launch
№ 069
LAUNCH
Source · Claude 开发者文档⚑ Product update

Anthropic's Cost Hack: Claude Fable 5 Advises, Claude Sonnet 5 Does the Heavy Lifting

The advisor model only chips in a few hundred words of guidance instead of doing the whole job — same quality, lower total cost. Currently a beta feature limited to the Claude API and AWS.

Claude APIMulti-Model Collaboration07-02 · 7 min read
Deep Dive
№ 068
DEEP DIVE
Source · Anthropic 官方文档⚑ Tutorial

Claude Fable 5's Official Prompting Guide: A Dozen Copy-Paste Fixes, From Over-Planning to Fake Progress Reports

Anthropic's official documentation shows you how to tune system prompts and engineering scaffolding for the new model — the same methods work for Claude Mythos 5 too.

Claude Fable 5Prompt Engineering07-02 · 9 min read
Deep Dive
🎧№ 067
DEEP DIVE
Source · Claude Blog⚑ Company PR

Claude Code's Official Playbook: 4 Levels of Agent Loops to Unattended

From manual confirmation to fully unattended, the Claude Code team lays out a 4-level loop taxonomy with practical guidance

Claude CodeAI Agent▶ Play07-02 · 8 min read
Free
Launch
№ 064
LAUNCH
Source · OASIS Devices 官网 / Digital Trends / Refresh Miami⚑ Product

OASIS Devices Unveils OASIS 1: A Smart Ring That Types From Silent Speech

Titanium build, $289 preorder, shipping around Christmas 2026 — the company previously shipped its first-gen touch-only ring.

Smart RingSilent Speech Input07-01 · 6 min read
Free
Launch
🎧№ 063
LAUNCH
Source · LongCat 官方博客 + The Decoder⚑ Product

Meituan Launches LongCat-2.0: 1.6-Trillion-Parameter Model Trained Entirely on Domestic Chips, No Nvidia GPUs

Trained on 50,000+ domestic AI chips and 35 trillion tokens; most benchmarks come from Meituan's own evaluation framework, and the weights aren't truly open for download yet

LLM ReleaseDomestic AI Chips▶ Play07-01 · 8 min read
Free
Launch
🎧№ 062
LAUNCH
Source · Anthropic⚑ Product

Anthropic Launches Claude Science: An AI Workbench for Scientists with 60+ Built-In Research Skills

Now in open beta. A coordinator agent marshals a team of expert agents to do the work, with a reviewer agent at the end dedicated to catching errors in citations and numbers — compute gets outsourced to AI, but raw data never leaves your local machine.

Claude ScienceResearch Agents▶ Play07-01 · 7 min read
Free
Launch
№ 061
LAUNCH
Source · 谷歌官方博客⚑ Product

Google Ships Nano Banana 2 Lite and Omni Flash Video Model: 4-Second Images, Fastest and Cheapest Yet

Images in just 4 seconds at about $0.034 per 1,000; the Omni Flash video model opens to developers the same day.

Nano BananaGemini Omni Flash07-01 · 5 min read
Research
🎧№ 060
RESEARCH
Source · Thinking Machines Lab⚑ Research

Bridgewater Built a Financial-Filtering Model with 84.7% Accuracy — and Open-Sourced the Method

Partnering with Thinking Machines, they fine-tuned an open-source model on expert-labeled data: 29.8% lower error rate than the best frontier model, at just 1/14 the inference cost

TinkerRL Fine-Tuning▶ Play07-01 · 7 min read
Free
Launch
№ 059
LAUNCH
Source · Anthropic 官网⚑ Product

Anthropic Launches Claude Sonnet 5: 40% Cheaper, Matches Opus 4.8 on Some Tasks

Official benchmarks show that at high-compute settings, it matches Opus 4.8 on some tasks — at just 60% of the standard price.

Claude Sonnet 5Anthropic07-01 · 7 min read
Deep Dive
🎧№ 058
DEEP DIVE
Source · Lucius / X⚑ Expert view

AI Replaced 700 Agents at Klarna. Then Quality Dropped.

Enterprise AI customer service has entered its consolidation era — Klarna and Alibaba's 2.56 million conversations both point to the same blind spot: cutting costs isn't the same as solving problems.

AI Customer ServiceEnterprise AI▶ Play06-30 · 9 min read
Deep Dive
№ 057
DEEP DIVE
Source · Big Technology AI Summit⚑ Expert view

OpenAI President: The Best Product Is No Product

Brockman confirms OpenAI is developing multiple hardware devices; Agent has only ~20 million users, while ChatGPT is nearing 1 billion

OpenAIAGI06-30 · 8 min read
Research
№ 056
RESEARCH
Source · Meta AI Blog⚑ Research

No Surgery Required: Meta's Non-Invasive Brain Reader Hits Nearly 8x the Accuracy

Just wear a helmet to decode brain-magnetic signals in real time — word accuracy jumps from 8% to 61%, with v1/v2 training code and datasets open-sourced simultaneously

Brain-Computer InterfaceMeta AI06-29 · 5 min read
Deep Dive
🎧№ 055
DEEP DIVE
Source · Loop Engineering (论文 PDF)⚑ Research

What Is Loop Engineering, Actually?

An Anthropic engineer's methodology for "loop engineering": instead of prompting AI one line at a time, design a self-running loop system

AI AgentEngineering Methodology▶ Play06-29 · 12 min read
Free
Launch
№ 054
LAUNCH
Source · GitHub⚑ Product

This Open-Source Tool Keeps Your Mac Awake Only as Long as the AI Is Running

One-click integration with 9 Agents including Claude Code — tasks keep running with the lid closed, and sleep control auto-releases within 50ms after the job stops.

macOSAI Agent06-29 · 4 min read
Free
Launch
№ 053
LAUNCH
Source · MarkTechPost / DeepSeek⚑ Product

One Extra Step, Way Faster: DSpark Boosts DeepSeek V4's Single-User Speed by 85%

60–85% faster on top of existing MTP-1 speculative decoding, by overlapping draft and verification in a pipelined execution (per DeepSeek's own benchmarks)

DeepSeekInference Acceleration06-29 · 5 min read
Deep Dive
№ 052
DEEP DIVE
Source · Every⚑ Company PR

The Fourth Step Everyone Skips Is What Actually Makes AI Engineering Work

Every runs five products with a one-person team — the core habit is one extra step after every feature ships: save the fix back into the system so AI automatically avoids the same trap next time.

AI CodingEngineering Methodology06-28 · 8 min read
Research
№ 051
RESEARCH
Source · METR⚑ Research

GPT-5.6 Sol's Cheating Rate Hits a Record High — Evaluators Say That's Reassuring

Three different versions of its capability score came out, and none of them can be trusted — but the visible cheating itself is evidence that safety monitoring works.

GPT-5.6AI Safety06-27 · 5 min read
Free
Launch
№ 050
LAUNCH
Source · OpenAI 官方博客⚑ Product

GPT-5.6 Just Launched — Except It Kind of Didn't

Three tiers at once — Sol, Terra, Luna — starting with a limited rollout to trusted partners (the list already filed with the US government), before wider access in a few weeks.

OpenAIGPT-5.606-27 · 7 min read
Research
№ 049
RESEARCH
Source · Wan Streamer⚑ Research

Wan Streamer: Real-Time AI That Listens, Watches, and Talks at Once

Model-side response ~200ms, end-to-end latency ~550ms; v0.1 caps out at 192p, and the demo is pre-recorded, not live

Multimodal LLMFull-Duplex Real-Time Interaction06-27 · 5 min read
Research
№ 043
RESEARCH
Source · IBM Newsroom⚑ Research

IBM Unveils World's First 0.7nm Chip, Packing Nearly 100 Billion Transistors on a Fingernail

Lab-verified as manufacturable; the +50% performance and +70% efficiency figures are projections versus 2nm, not measured results

IBMSemiconductors06-26 · 6 min read