Deep Dive · XiaoHu Brief

Princeton professor at ICML on AI and work: how should individuals adapt?

In ~24 months, model capability rose sharply; SAGE’s composite reliability metric rose only 5–10 percentage points
60-second takeaway
  • Princeton CS professor Arvind Narayanan’s ICML 2026 keynote in Seoul, “What will be left for us to work on?”, builds on the “AI as Normal Technology” framework with Sayash Kapoor.
  • SAGE lab: over ~24 months, frontier models jumped in capability; reliability (consistency / robustness / calibration / operational safety) rose only 5–10pp.
  • In software’s decide–execute–deliver stack, AI mainly compresses execute (~1/3 of effort); decide and deliver are not compressed—and may expand.
  • ATM, radiology, translation, software tools: automation rarely cuts jobs 1:1; software employment grew ~10,000× through many ~10× tool leaps.
  • He splits RSI / human-level AI / economically transformative AI / superintelligence into four non-entailing dimensions.
  • Personal adaptation: raise the ceiling, don’t only ride the floor; balance productivity / growth / control; refuse black boxes; master first, then amplify; reinvest ~10 hours/week into skills.
1Opening

Two narratives: the AI field is anxious about its own jobs

At ICML 2026 in Seoul, Princeton CS professor Arvind Narayanan gave a keynote titled “What will be left for us to work on?”

He faces the anxiety head-on: as AI can do more of what we do, how should we prepare?
🎯
Why it matters: He runs SAGE on AI agent evaluation, co-authored the long essay AI as Normal Technology (~15k words, becoming a book). The talk ties the framework, self-built reliability data, and frontline software observations together.
What will be left for us to work on?
Title slide. Source: Arvind Narayanan · ICML 2026 keynote slides

He splits the path into two practical camps—not pure philosophy:

Narrative 1 · Replacement
In a few years AI replaces almost everything we do today. Rational move: accumulate wealth before skills devalue. Some in Silicon Valley take this path; the “permanent underclass” meme lives outside the lab too.
Narrative 2 · Amplification (his camp)
AI will greatly amplify human potential. Now is the best time to build complementary skills—agency, taste, judgment—and the scaffolding around them.
Two narratives
Replacement vs amplification. Source: keynote slides

If you bet on replacement and amplification wins, you may miss the best window in history to build superpowers. The world is watching how the AI community responds; rolling over and accepting “AI will do the work” may fuel a sharper political backlash.

2Framework

Four stages of AI impact—the slowest has barely begun

“Normal” in AI as Normal Technology does not mean AI is a hammer or a toothbrush. They treat it as industrial-revolution scale, as a causal model of how capability becomes economic and social impact.

Classic diffusion: invention → innovation (appliances) → adoption. They expand it into four stages, with software as the example:

① Methods / capabilityModels improve fast
② Products / appsCoding agents, etc.
③ Early adoptionVibe coding → agentic engineering
④ AdaptationOrg redesign over decades
Four stages
Diffusion applied to AI. Source: keynote slides

Stage four is slowest. Even in software—an early coding-agent adopter—true organizational redesign has barely started. Speculation: if agents can ship huge, secure codebases, one-size-fits-all software for billions makes less sense; software becomes extremely personalized, and even “software company” as a form may be renegotiated. That is human and organizational change—historically measured in decades.

3History

Electricity in factories: drop-in replacement never worked

Pre-electricity factories used one giant steam engine and mechanical transmission. Owners first tried swapping boilers for generators—“more efficient drop-in.” It failed.

Electricity transformation
Gains came from reorganizing, not swapping boilers. Source: keynote slides
Failed path · Drop-in
New power source, same layout, same roles, same institutions. Efficiency hopes collapsed.
Working path · ~40 years of redesign
Power to any station → assembly-line layout → training, hiring/firing, labor law redesigned together.

Today agents are sold as drop-in human replacements. Electricity’s lesson: the payoff is reorganizing work, not cloning a person with a machine. That was not the utility’s job—and AI’s org redesign is not only AI companies’ job. In their four stages, this is the slowest; it has barely opened.

4Evidence

Measured: capability rose; reliability lagged

There is a huge gap between what occupations could use AI for and what they actually use. Adoption lag is one reason; deployers may also feel walls beyond leaderboard scores sooner than vendors admit.

Reliability is the #1 cited concern. SAGE clusters ~10–12 metrics into four dimensions:

Consistency
Does “70% accuracy” mean 70% of tasks always work—or every task fails 30% of the time? Deployment meaning differs; many agent benchmarks do not separate them.
Robustness
Does performance collapse after small environment shifts?
Calibration
Can the agent tell whether it just succeeded?
Operational safety
On failure, can it recover—or delete production data?
~24 mo
Window of large capability jumps at three frontier labs
5–10pp
Rise in SAGE’s composite reliability metric in the same window
4 dims
Consistency · robustness · calibration · ops safety
2 of 3
General + high-stakes + fully autonomous: pick two (his inference)
Capability vs reliability (schematic, not raw axes) Capability / accuracy · steep Reliability · ~+5–10pp ~24 months · 3 frontier labs (SAGE narrative)
Capability and reliability rise at different speeds—the talk’s hardest empirical claim.
Capability-reliability gap
Capability–reliability gap. Source: keynote slides · SAGE
Inference

General + high-stakes + fully autonomous looks like pick-two. Collaboration agents will keep outperforming full automation agents; scaffolding and post-training should differ—not one “more automatic is better” story.

5Model 1

Software engineering: writing code was never the bottleneck

Papers said this by 2019. The last year of blogs “rediscovered” it: coding agents sped the middle layer; the whole job did not shrink in proportion.

D
Decide
Requirements, design, planning—not clearly compressed by AI.
E
Execute (compressed)
Coding and debugging—compressed by agents; was only ~1/3 of the work.
D
Deliver
Owning releases, integration, maintenance, testing—not compressed; may grow as execute speeds up.
decide-execute-deliver sandwich
The sandwich: AI mostly bites the middle. Source: keynote slides
Crane operator

Machines do cognitive heavy lifting; humans stay in control. The job becomes operating the machine, not hauling every brick by hand. Cranes did not erase job sites—they rewrote who does what.

Crane metaphor
Useful metaphor: crane operator. Source: keynote slides
6History

Lump-of-labor fallacy: automation rarely cuts jobs 1:1

The fallacy: work is a fixed pie; AI takes a slice and jobs permanently vanish. History often shows demand and job structure changing with efficiency.

Lump of labor
Lump-of-labor fallacy and historical cases. Source: keynote slides
ATM × bank tellers
ATMs enabled more branches; tellers shifted to non-ATM work—not one ATM, one fewer teller forever.
Radiology × Hinton’s prophecy
Hinton roughly predicted radiologists “basically extinct in five years.” Employment grew; radiologists adopt AI rather than blanket-reject it.
Lawsuits × lawyers
AI makes filing easier → more litigation and lawyer work (bad for society maybe; not obviously bad for lawyer employment).
MT × translators
Near-human MT for ~a decade; translator employment roughly stable and projected stable—no hard ceiling on content × language pairs.

Software itself: many ~10× tool leaps from machine code onward; employment rose ~10,000× because code demand grew faster. Not “nobody ever loses a job”—just “automation rate ≠ unemployment rate.”

7Model 2

If RSI arrives: four dimensions that get smushed together

Labs race toward recursive self-improvement (RSI). He takes it seriously—without equating a lab milestone with “humans instantly have nothing left to do.”

Hawaii problem
The Hawaii problem in AGI talk. Source: keynote slides
Hawaii problem

Early explorers could name a whole archipelago “Hawaii.” Up close, if you still do not name islands, you get lost about where to sail. RSI, human-level AI, economically transformative AI, and superintelligence are often chained as automatic dominos—he wants them unbundled.

Four dimensions
Four dimensions of advanced AI—no automatic entailment. Source: keynote slides

Curing cancer is often gated by thousand-patient, multi-year trials—external constraints more compute cannot delete. Mapping “model milestone” straight to “society has no work left” skips those walls.

8Creativity

Why AI creativity lags—and open-world evaluation

Perception representations are strong; representations that support creativity and high-level reasoning still lag, in his view. Hypotheses from cognitive science and practice:

Compositional reps
Chollet et al.: human creativity from extreme composition of few “meaning atoms”; working-memory limits force efficiency. AI’s memory/retrieval strengths make fair comparison hard.
Updating reps while thinking
Humans improve problem representations mid-thought (“sleep on it”). Today’s systems largely cannot.
Continual learning gap
Field mostly fights catastrophic forgetting; little work on steadily raising representation quality.
Open-world eval
No answer key; real open problems judged by experts—less exam hacking / data contamination.
Creativity paradox
AI creativity paradox. Source: keynote slides

SAGE open-world evaluation: give agents a few thousand dollars plus a real ML problem a human expert spent months on (paper not yet on arXiv), then have those experts grade outputs. They also ran “ship an iOS app to the App Store” style tasks and are recruiting senior researchers to expand.

9Personal adaptation

How should individuals adapt in this wave?

After the framework, Narayanan’s “Personal reflections on adapting to AI” is not a universal playbook—he shows how he is surviving the capability treadmill in his own research workflows.

Personal reflections
Personal reflections on adapting to AI. Source: keynote slides

Pick a stance: replacement vs amplification

The opening narratives become life configurations:

Bet on quick replacement
Race to wealth before skills devalue. If amplification wins, you may miss the best skill-building window in history.
Bet on amplification
Build complementary skills now—agency, taste, judgment. That is his camp.

Floor and ceiling: where reinvested time goes

Floor
What AI can do alone. The floor rises; industry stories obsess over this line.
Ceiling
What you can take on with AI—ambitious projects that were impossible before. The ceiling does not rise automatically; only if you push it.

His practice: when AI buys big productivity, reinvest the surplus into long-term growth and complementary skills/workflows—about 10 hours/week learning and experimenting.

A harsh self-check

“If I don’t feel exhausted at the end of the day, I’ve done something wrong. I’ve offloaded too much to AI—sacrificing long-term growth for short-term productivity.”

Three-legged stool: productivity · growth · control

ProductivityShip the work
+
GrowthSkills & workflows compound
+
ControlHuman keeps the wheel

Only productivity → crushed when the floor rises. Only growth → you do not ship. Both without control → a button-clicker over time.

Two heuristics to keep control

① Resist the black box
Vendors want one prompt and a full agent run. He calls that a trap that slowly surrenders control.
② Avoid the dependence spiral
Tempting to hand AI tasks you are not yet expert at (learning is hard). Residual skill then erodes. Better long run: master first, then amplify.
Back to the crane

Machines lift; humans stay in the cab. The job becomes operating, understanding, and controlling the machine—not being a brick on the schedule. Adaptation means not climbing out of the cab.

Crane operator
Jobs redefined as operating and controlling the machine. Source: keynote slides
10Vision

Closing vision: human–AI co-superintelligence

Economically transformative AI is already starting—not via an AGI switch, but via slow variables: reliability, integration, tacit knowledge, regulation. He rejects geopolitics as “whoever hits a capability milestone first takes all economic value.”

On superintelligence: tasks have ceilings; human intelligence leans on learning and tools, so AI is another tool and the contest is more “AI-augmented humans” vs “AI alone”; treating “AI owns companies and hires/fires” as default and hoping alignment alone saves us is “less pro-safety than anti-safety.”

co-superintelligence
Vision: co-superintelligence. Source: keynote slides

Computers were “bicycles for the mind”; he calls AI a “crane for the mind.” The learning curve is steep—a treadmill—but co-superintelligence is a fight worth fighting: not abandoning work, but redefining it as higher-ceiling collaboration with AI.

Takeaways
Takeaways slide. Source: keynote slides
The capability floor rises on its own; the ceiling you must push. Adaptation is not dumping all work on AI—it is reinvesting saved time into complementary skill while staying in the driver’s seat. From Arvind Narayanan · ICML 2026 keynote
Based on Arvind Narayanan’s ICML 2026 keynote “What will be left for us to work on?” and the AI as Normal Technology write-up (2026-07-14). Figures such as 5–10pp reliability, ~24-month window, execute ~1/3, and ~10,000× software employment are as stated in the talk; SAGE metrics are lab-built. Slides: cs.princeton.edu/~arvindn/talks/icml-2026-annotated-slides/ · Source: normaltech.ai