Library / Research

Research

35 explainers · newest first

Free
Research
№ 1101
RESEARCH
Source · BleepingComputer⚑ Research

LayerX finds a new visual fraud attack that tricks AI browser assistants

A custom font and a few lines of CSS are enough—no JavaScript, and no browser vulnerability required.

AI securitybrowser assistants07-30 · 10 min read
Research
🎧№ 1099
RESEARCH
Source · OpenAI⚑ Tutorial

Two API settings, not a new model: OpenAI reports GPT-5.6 Sol jumping from 13.3% to 38.3% on ARC-AGI-3

The model stayed the same; the harness did not—retained private reasoning plus compaction instead of deletion cut output tokens per game to about one-sixth.

ARC-AGI-3harness▶ Play07-30 · 9 min read
Free
Research
№ 1097
RESEARCH
Source · Anthropic⚑ Research

Claude Mythos finds pure mathematical flaws in encryption schemes, matching elite cryptographers

Neither result threatens anything in production: HAWK is not deployed yet, and the AES work hit only a seven-round reduced version, not the full ten-round standard.

CryptographyAI research07-29 · 13 min read
Free
Research
🎧№ 1092
RESEARCH
Source · OpenAI Economic Research⚑ Research

OpenAI analyzed 800,000 work-related messages: most of what people use AI for is not what their job was supposed to be

The exact same dataset produces task crossover rates of 43.5% and 65%–82%, differing only in how the denominator is sliced.

OpenAILabor economics▶ Play07-28 · 12 min read
Free
Research
🎧№ 1089
RESEARCH
Source · Kimi K3 技术报告⚑ Research

Kimi K3 technical report: How three architectural changes boosted compute efficiency 2.5x

Published 11 days after launch, the 47-page paper focuses on efficiency, delivering 2.5 times the performance of K2 on the same compute budget.

Kimi K3Moonshot AI▶ Play07-28 · 16 min read
Free
Research
№ 1085
RESEARCH
Source · Google Blog / AI & Economy ATLAS v1.0⚑ Research

Google combed through 14.65 million Gemini chats — 86% weren't about work

The median occupation has AI touching just one-fifth of its tasks, 29% of jobs show zero AI use at all, and even in cognitive work, AI carries a task start to finish only 6.5% of the time.

GoogleAI and the Economy07-27 · 14 min read
Research
№ 1071
RESEARCH
Source · Databricks⚑ Company PR

Databricks benchmarks AI coding agents on millions of lines of real code

Swapping the harness around the same model can double the cost, and open-source GLM 5.2 matches Opus 4.8 for 30% less per task — on a benchmark built from Databricks' own merged pull requests, so none of it is searchable online.

Coding agentsBenchmarks07-22 · 9 min read
Research
№ 1064
RESEARCH
Source · Memories.ai Research⚑ Research

Memories.ai's O-MARC makes audio-visual AI faster, cheaper, and more accurate

The release ships with a companion benchmark that strips out the audio track and re-runs the test, filtering out questions models can already answer by sight alone.

Video understandingMultimodal compression07-21 · 12 min read
Research
🎧№ 1058
RESEARCH
Source · New Scientist⚑ Expert view

Study finds a sweet spot for AI-assisted writing — not too much, not too little

Three separate trials of the same study all landed on the middle ground — and the columnist who tried ChatGPT on a movie synopsis says he'd still rather write it himself.

AI WritingCreativity▶ Play07-20 · 7 min read
Research
🎧№ 1053
RESEARCH
Source · arXiv⚑ Research

The model you picked through a router might be swapped — one random number proves it

The fingerprint distance between two samples of the same model has a median of 0.140. One API marketed as a proprietary in-house flagship scores 0.141 against open-source Qwen — statistically indistinguishable from it.

Model fingerprintingAPI auditing▶ Play07-20 · 11 min read
Research
🎧№ 1042
RESEARCH
Source · Impossible Research⚑ Research

A single wrapper pushes Opus 4.8 and Fable 5 to 99% on ARC-AGI-3

A harness called Schema has models turn each game's rules into a runnable, verified program before making a move. Across 25 public rounds it self-reported 98.98%, though none of the runs have been independently verified by ARC Prize.

ARC-AGI-3World Model▶ Play07-17 · 12 min read
Free
Research
№ 1035
RESEARCH
Source · Design Arena

GPT-5.6 Sol tops front-end design leaderboard by dodging AI's usual clichés

Design Arena combed through a thousand of the model's web pages and found a gap in the design space — right where clichés like purple gradients usually cluster.

GPT-5.6 SolAI web design07-16 · 6 min read
Free
Research
№ 1027
RESEARCH
Source · MIT News / Science⚑ Research

MIT Builds a Robotic Bird That Swims and Flies Without Paddling Feet

A 250-gram robot crosses water and air with the same flexible wings, taking off from a lake at a 70° pitch after 8–10 wingbeats

MITBio-inspired robotics07-15 · 8 min read
Free
Research
🎧№ 1020
RESEARCH
Source · PrismML⚑ Research

PrismML Stuffs a 27B Model into Your iPhone — With Barely Any IQ Drop

It squeezes a ~54GB 27B model down to about 3.9–5.9GB: it runs on-device, and average scores still hold about nine-tenths. Core value first; technical detail and the fine print come later.

Bonsai 27BOn-device AI▶ Play07-15 · 12 min
Free
Research
🎧№ 1014
RESEARCH
Source · Anthropic⚑ Research

Anthropic Analyzed 300,000 Conversations: Claude's Values Shift by Language

Across 3 models and 20 languages: English is the most cautious and in-depth, Russian the most exacting, Hindi the warmest, and Chinese sits closest to the global average

Anthropic ResearchClaude Values▶ Play07-14 · 9 min read
Free
Research
🎧№ 1013
RESEARCH
Source · Sakana AI / Nature Communications⚑ Research

Sakana AI's Brainless 3D Cell Bricks Self-Organize With Biology-Like Swarm Intelligence

With no central brain in charge, nearly 200 simple smart cubes figure out what shape they've formed just by talking to their neighbors—and can even sense where to "regrow" after damage. The self-recognition part already works on physical bricks; damage localization and regeneration still happen mostly in simulation.

Sakana AISwarm Intelligence▶ Play07-14 · 9 min read
Research
🎧№ 1010
RESEARCH
Source · arXiv · Meta AI⚑ Research

Meta AI's Proactive Memory Agent Teaches Models When Not to Remind You, Boosting Terminal-Bench Accuracy by 8.3 Points

The paper claims the code is open-sourced — but the repo turns out to be empty, without a single commit ever pushed.

AI AgentLong-Task Memory▶ Play07-13 · 9 min read
Research
🎧№ 1008
RESEARCH
Source · Anthropic 官方博客⚑ Research

Claude Cowork's Top Use Case Is Office Admin (33.4%), Not Coding (8.7%): Interface Shapes How AI Gets Used

Based on 1.2M+ conversations across 600,000+ organizations: content creation ranks second at 16.4%, together accounting for nearly half of all usage.

Claude CoworkAnthropic▶ Play07-13 · 7 min read
Free
Research
🎧№ 1006
RESEARCH
Source · VentureBeat⚑ Research

The Package AI Told You to Install Doesn't Exist — Hackers Already Named Their Malware After It

576K samples, 16 models tested: 19.7% of AI-recommended packages are hallucinations, and 43% keep generating the same fake name.

AI SecuritySoftware Supply Chain▶ Play07-12 · 9 min read
Research
🎧№ 1001
RESEARCH
Source · EPFL 项目主页⚑ Research

What Does Each Brain Region Like to Watch? EPFL Evolved AI-Generated Videos to Find Out

All results are computer simulation predictions from a brain "digital twin" model, not yet validated with real human brain imaging.

NeuroscienceEvolutionary Algorithm▶ Play07-11 · 9 min read
Research
🎧№ 112
RESEARCH
Source · Google Research⚑ Research

Google Research Unveils SensorFM: Trained on a Trillion Minutes of Wearable Data, It Wins 33 of 35 Health Tasks

Pretrained on 2 billion hours of wearable data from 5 million people, a frozen encoder with just a linear head beats supervised baselines on 34 of 35 health tasks.

SensorFMWearable Health Data▶ Play07-10 · 8 min read
Research
🎧№ 104
RESEARCH
Source · arXiv / Google DeepMind⚑ Research

Gemma 4 Technical Report: How a Small Open Model Takes On Large Ones with Reasoning and Memory Efficiency

Not a feature list — three engineering tracks turned at once: reasoning lifts intelligence, an efficiency stack cuts cost, native multimodality expands input.

Gemma 4Open-Source LLMs▶ Play07-09 · 12 min read
Research
🎧№ 103
RESEARCH
Source · LangChain Blog⚑ Research

LangChain Tunes the Harness, Not the Model — Nemotron 3 Ultra Closes In on Opus 4.8 at 1/10 the Cost

Three levers — system prompt, tool descriptions, middleware — push the Deep Agents suite from a typical ~0.80 to 0.84, topping out at 0.86 against Opus's 0.87.

LangChainNemotron▶ Play07-09 · 9 min read
Research
🎧№ 102
RESEARCH
Source · Lil'Log⚑ Research

AI Self-Improvement Starts Outside the Model: Lilian Weng on Harness Engineering

Former OpenAI safety lead surveys nearly 30 papers: from prompt tweaks to self-modifying code, DGM pushed coding ability from 20% to 50%.

Harness EngineeringRecursive Self-Improvement▶ Play07-09 · 13 min read
Research
🎧№ 096
RESEARCH
Source · Liquid AI⚑ Research

Liquid AI's Antidoom Fixes Reasoning Models' Doom Loops With a Single Token — Now Open-Sourced

By fine-tuning only the single token where the doom loop begins, both models' loop rates drop to around 1%

Liquid AIReasoning Model Doom Loops▶ Play07-08 · 8 min read
Free
Research
🎧№ 088
RESEARCH
Source · Anthropic⚑ Research

Anthropic Discovers a Brain-Like 'Inner Workspace' in Claude — It Evolved, Wasn't Designed

It makes up less than 10% of the model — remove it and Claude can still talk, but its reasoning collapses to zero. Anthropic is already using it to catch fabricated data and spot when Claude senses it's being tested.

Interpretability ResearchClaude▶ Play07-07 · 8 min read
Free
Research
🎧№ 087
RESEARCH
Source · iTextbooks'26 论文⚑ Research

Dartmouth Put AI Grading to the Test: Students Called It Rigid — But Users Scored Higher

In a 151-student trial, short-answer questions moved scores more than multiple choice, while almost no one touched the AI help sidebar

AI in EdTechFormative Assessment▶ Play07-06 · 8 min read
Free
Research
🎧№ 083
RESEARCH
Source · Seldo.com(Laurie Voss)⚑ Research

AI Is Torching Junior Dev Jobs: Coding Is Becoming a Basic Skill, Not a Career

US developer employment among 22-to-25-year-olds has fallen 19% in three years, even as new GitHub sign-ups hit their fastest growth ever.

AI Job DisruptionJunior Developers▶ Play07-05 · 9 min read
Research
🎧№ 082
RESEARCH
Source · arXiv · NVIDIA Research⚑ Research

NVIDIA Research Unveils HORIZON: Unattended Agents Push Full RTL Chip Design Benchmark to 100% Pass Rate

The paper is the first to run an agent through an entire RTL benchmark suite fully unattended—most tasks clear in two or three rounds, but the hardest one takes 82 iterations

NVIDIA ResearchAgentic Workflows▶ Play07-05 · 12 min read
Research
🎧№ 081
RESEARCH
Source · Anthropic⚑ Research

Anthropic Analyzed 400K Claude Code Sessions: Expertise Beats Coding Skill

A 7-month analysis of sessions from 235,000 users: verified experts succeed at nearly double the rate of novices — yet the top 10 professions differ by no more than 7 percentage points.

Claude CodeAnthropic Research▶ Play07-05 · 7 min read
Research
🎧№ 060
RESEARCH
Source · Thinking Machines Lab⚑ Research

Bridgewater Built a Financial-Filtering Model with 84.7% Accuracy — and Open-Sourced the Method

Partnering with Thinking Machines, they fine-tuned an open-source model on expert-labeled data: 29.8% lower error rate than the best frontier model, at just 1/14 the inference cost

TinkerRL Fine-Tuning▶ Play07-01 · 7 min read
Research
№ 056
RESEARCH
Source · Meta AI Blog⚑ Research

No Surgery Required: Meta's Non-Invasive Brain Reader Hits Nearly 8x the Accuracy

Just wear a helmet to decode brain-magnetic signals in real time — word accuracy jumps from 8% to 61%, with v1/v2 training code and datasets open-sourced simultaneously

Brain-Computer InterfaceMeta AI06-29 · 5 min read
Research
№ 051
RESEARCH
Source · METR⚑ Research

GPT-5.6 Sol's Cheating Rate Hits a Record High — Evaluators Say That's Reassuring

Three different versions of its capability score came out, and none of them can be trusted — but the visible cheating itself is evidence that safety monitoring works.

GPT-5.6AI Safety06-27 · 5 min read
Research
№ 049
RESEARCH
Source · Wan Streamer⚑ Research

Wan Streamer: Real-Time AI That Listens, Watches, and Talks at Once

Model-side response ~200ms, end-to-end latency ~550ms; v0.1 caps out at 192p, and the demo is pre-recorded, not live

Multimodal LLMFull-Duplex Real-Time Interaction06-27 · 5 min read
Research
№ 043
RESEARCH
Source · IBM Newsroom⚑ Research

IBM Unveils World's First 0.7nm Chip, Packing Nearly 100 Billion Transistors on a Fingernail

Lab-verified as manufacturable; the +50% performance and +70% efficiency figures are projections versus 2nm, not measured results

IBMSemiconductors06-26 · 6 min read