A custom font and a few lines of CSS are enough—no JavaScript, and no browser vulnerability required.
The model stayed the same; the harness did not—retained private reasoning plus compaction instead of deletion cut output tokens per game to about one-sixth.
Neither result threatens anything in production: HAWK is not deployed yet, and the AES work hit only a seven-round reduced version, not the full ten-round standard.
The exact same dataset produces task crossover rates of 43.5% and 65%–82%, differing only in how the denominator is sliced.
Published 11 days after launch, the 47-page paper focuses on efficiency, delivering 2.5 times the performance of K2 on the same compute budget.
The median occupation has AI touching just one-fifth of its tasks, 29% of jobs show zero AI use at all, and even in cognitive work, AI carries a task start to finish only 6.5% of the time.
Swapping the harness around the same model can double the cost, and open-source GLM 5.2 matches Opus 4.8 for 30% less per task — on a benchmark built from Databricks' own merged pull requests, so none of it is searchable online.
The release ships with a companion benchmark that strips out the audio track and re-runs the test, filtering out questions models can already answer by sight alone.
Three separate trials of the same study all landed on the middle ground — and the columnist who tried ChatGPT on a movie synopsis says he'd still rather write it himself.
The fingerprint distance between two samples of the same model has a median of 0.140. One API marketed as a proprietary in-house flagship scores 0.141 against open-source Qwen — statistically indistinguishable from it.
A harness called Schema has models turn each game's rules into a runnable, verified program before making a move. Across 25 public rounds it self-reported 98.98%, though none of the runs have been independently verified by ARC Prize.
Design Arena combed through a thousand of the model's web pages and found a gap in the design space — right where clichés like purple gradients usually cluster.
A 250-gram robot crosses water and air with the same flexible wings, taking off from a lake at a 70° pitch after 8–10 wingbeats
It squeezes a ~54GB 27B model down to about 3.9–5.9GB: it runs on-device, and average scores still hold about nine-tenths. Core value first; technical detail and the fine print come later.
Across 3 models and 20 languages: English is the most cautious and in-depth, Russian the most exacting, Hindi the warmest, and Chinese sits closest to the global average
With no central brain in charge, nearly 200 simple smart cubes figure out what shape they've formed just by talking to their neighbors—and can even sense where to "regrow" after damage. The self-recognition part already works on physical bricks; damage localization and regeneration still happen mostly in simulation.
The paper claims the code is open-sourced — but the repo turns out to be empty, without a single commit ever pushed.
Based on 1.2M+ conversations across 600,000+ organizations: content creation ranks second at 16.4%, together accounting for nearly half of all usage.
576K samples, 16 models tested: 19.7% of AI-recommended packages are hallucinations, and 43% keep generating the same fake name.
All results are computer simulation predictions from a brain "digital twin" model, not yet validated with real human brain imaging.
Pretrained on 2 billion hours of wearable data from 5 million people, a frozen encoder with just a linear head beats supervised baselines on 34 of 35 health tasks.
Not a feature list — three engineering tracks turned at once: reasoning lifts intelligence, an efficiency stack cuts cost, native multimodality expands input.
Three levers — system prompt, tool descriptions, middleware — push the Deep Agents suite from a typical ~0.80 to 0.84, topping out at 0.86 against Opus's 0.87.
Former OpenAI safety lead surveys nearly 30 papers: from prompt tweaks to self-modifying code, DGM pushed coding ability from 20% to 50%.
By fine-tuning only the single token where the doom loop begins, both models' loop rates drop to around 1%
It makes up less than 10% of the model — remove it and Claude can still talk, but its reasoning collapses to zero. Anthropic is already using it to catch fabricated data and spot when Claude senses it's being tested.
In a 151-student trial, short-answer questions moved scores more than multiple choice, while almost no one touched the AI help sidebar
US developer employment among 22-to-25-year-olds has fallen 19% in three years, even as new GitHub sign-ups hit their fastest growth ever.
The paper is the first to run an agent through an entire RTL benchmark suite fully unattended—most tasks clear in two or three rounds, but the hardest one takes 82 iterations
A 7-month analysis of sessions from 235,000 users: verified experts succeed at nearly double the rate of novices — yet the top 10 professions differ by no more than 7 percentage points.
Partnering with Thinking Machines, they fine-tuned an open-source model on expert-labeled data: 29.8% lower error rate than the best frontier model, at just 1/14 the inference cost
Just wear a helmet to decode brain-magnetic signals in real time — word accuracy jumps from 8% to 61%, with v1/v2 training code and datasets open-sourced simultaneously
Three different versions of its capability score came out, and none of them can be trusted — but the visible cheating itself is evidence that safety monitoring works.
Model-side response ~200ms, end-to-end latency ~550ms; v0.1 caps out at 192p, and the demo is pre-recorded, not live
Lab-verified as manufacturable; the +50% performance and +70% efficiency figures are projections versus 2nm, not measured results