Predictable Generative UI
Our interfaces will soon generate themselves, unique to each person and moment. To stay usable, they must stay predictable. Eight requirements. Thirty years of evidence. Every claim cited.
§0 · Foreword
Four pixels cost ten percent
Early in my time leading design teams at scale, a team adjacent to mine rounded the corners of listing cards by four pixels, on a surface serving hundreds of millions of people. It caused a 10% regression. The change wasn't wrong in isolation; it was wrong at scale, inside an interface people had already learned. That experiment wasn't about predictability. It taught me something adjacent and more fundamental: once people have learned a surface, even tiny changes to it carry enormous consequences.
Hold that thought against what the industry is building next: interfaces that regenerate themselves (layout, controls, hierarchy) every time you open them. I've believed for a long time that a system earns trust by being predictable in how it behaves. Generative UI is the hardest test that belief has ever faced.
So I went back through thirty years of adaptive-interface research and the last two years of generative-UI evidence, and I wrote down what I found. I wrote it as a specification, not an essay, because this moment doesn't need another opinion about AI. It needs constraints.
ProvenanceAltaffer, “Three convictions,” 2026·Full references · Annex A
§1 · Scope
What this document covers
This document applies to any interface assembled by a model at runtime: layouts generated per prompt, dashboards composed per user, tool panels that appear because an AI decided you need them. It is written for the people shipping those systems: designers, engineers, and the leaders deciding how much freedom the model gets.
It specifies eight requirements (REQ‑01 through REQ‑08) that bound what a generative system may change, how fast, and with how much explanation.
Two things this document is not. It is not an argument against generative UI; the promise is real, and I want it kept. And it is not a style guide; it says nothing about how generated interfaces should look.
It is about one property only: whether the person using the interface can still locate themselves, predict what happens next, and feel in control while the interface changes underneath them.
§2 · The promise
Every user gets exactly the interface they need
Nielsen Norman Group defines generative UI as an interface “dynamically generated in real time by AI to fit the user's needs and context.” Not one design for a billion people: a design for you, for this task, for this moment. Designers stop crafting screens and start defining outcomes.
The industry has committed real money to this. Google generates entire interfaces per prompt in Gemini. OpenAI and Anthropic ship interactive apps inside their assistants. Vercel and Thesys sell the infrastructure. A €10M German consortium (Audi, MAN, Valeo) is building it into vehicle interiors.
It's a real promise. I'm not here to argue with it. I'm here to keep it from failing.
EvidenceMoran & Gibbons, NN/g 2024·Google Research 2025·OpenAI 2025·Anthropic 2026·Fraunhofer SALSA
§2.2 · The record: twelve months of field data
The industry ran the experiment. Watch the correction.
| Date | Event | Reading |
|---|---|---|
| 2025‑10 | OpenAI ships the Apps SDK: developer-authored, reviewed components rendered in chat. The model orchestrates; the UI holds still. | Structure, not generation |
| 2025‑11 | Google ships fully generative UI in Gemini. Raters prefer it to top websites 90% of the time, yet expert human designs still beat it, 56% to 43%. Jakob Nielsen calls it an “interaction synthesizer.” | The maximalist bet: peak hype |
| 2026‑01 | Anthropic ships MCP Apps: sandboxed, developer-authored UI components in Claude. OpenAI adopts the same open standard the same week. | Both labs converge on stable components |
| 2026‑04 | CHI dedicates a workshop to whether generated UIs can sustain the stable behavior users' mental models depend on. | The field names the open question |
| 2026‑07 | Nielsen scores his own generative-UI predictions at 47% and lands on “a stable skeleton and generated flesh”: navigation, identity, and undo stay fixed, with local panels generated on demand. | The optimist arrives at the constraint position |
I wrote the first revision of this document in March 2026. The field converged on it by July.
EvidenceOpenAI Apps SDK·Leviathan et al., Google Research 2025·Nielsen, Nov 2025·MCP Apps spec, Jan 2026·CHI 2026 workshop·Nielsen, July 2026
§3 · Prior art: we have run this experiment before
Prior art
“In the past, adaptive user interfaces have widely failed user acceptance.”
Fraunhofer IOSB, Human-AI Interaction, the institute now building generative UIs for Audi and MAN
Adaptive interfaces (systems that rearrange themselves to fit you) are not a new idea. They have three decades of results, and the results are consistent.
Findlater and McGrenere put static, adaptive, and adaptable menus head to head: the static menu was faster than the adaptive one. Lavie and Meyer tested adaptivity in vehicles: it helped in routine situations, and in unfamiliar ones it raised cognitive workload and hurt performance. The one early success, Sears and Shneiderman's split menus, moved a few frequent items, slowly, and left everything else alone.
Users reject interfaces that change unpredictably, even when the changes are objectively better. That sentence should stop everyone shipping generative UI mid-sprint.
We never lacked the technology to personalize interfaces. We lacked a reason users would tolerate it.
EvidenceFraunhofer IOSB, GenUIn·Findlater & McGrenere, CHI 2004·Lavie & Meyer, IJHCS 2010·Sears & Shneiderman, ToCHI 1994
§3.2 · The controlled result
Accuracy buys tolerance. It doesn't buy forgiveness.
The cleanest evidence comes from Gajos, Everitt, Tan, Czerwinski, and Weld at CHI 2008. They isolated the two properties everyone argues about: predictability (can you anticipate what the interface will do) and accuracy (does the adaptation actually match what you need).
Both improved satisfaction. But accuracy had the stronger effect on performance and utilization, stronger than the authors themselves expected. Read carefully, that's not a license to generate freely. It's a warning with two edges: when the system adapts correctly, users tolerate some surprise. When it adapts wrongly (and every generative system sometimes will), unpredictability makes the failure doubly punishing. You guessed wrong, and I can't find anything.
Plan for the miss, not the hit. Predictability is what failure costs depend on.
§3.3 · Hidden costs: the ones that don't show up in task time
Even when adaptation works, it takes something
It narrows what you know. Findlater and McGrenere found that higher-accuracy adaptive menus improved performance while reducing users' awareness of the features the adaptation hid. The interface got faster at today's task by quietly shrinking the user's understanding of the system.
And the obvious fix backfires. Ease changes in slowly? At CHI 2018, Brock and Quigley tested interfaces that shifted with the user's proximity: the moving interface produced change-blindness errors at a rate 13.7 percentage points higher than a static one, and gradual changes induced more errors than instant ones. Slow-crept change slips under attention entirely.
Change has to be both noticeable and expected. Slipping it past the user is not a courtesy. It's a defect.
EvidenceFindlater & McGrenere, CHI 2008·Brock & Quigley et al., CHI 2018
§4 · Mechanism: why the rejection is rational
Your brain bills every change as a loss
None of this is users being stubborn. It's accounting. Kahneman and Tversky showed that losses loom roughly twice as large as equivalent gains, and Samuelson and Zeckhauser showed we defend the current state even against objectively better alternatives. A learned layout is a possession; Kahneman, Knetsch, and Thaler called it the endowment effect. Kim and Kankanhalli measured the same force in enterprise software: users resist beneficial systems because switching costs are paid now and benefits are paid later.
Run the math for a generated layout. A rearrangement that is 10% more efficient still lands as a loss, because the user pays the reorientation cost immediately and collects the efficiency slowly, against a roughly two-to-one penalty on the loss side.
A generated change must be about twice as good just to feel neutral. Most aren't.
EvidenceKahneman & Tversky 1979·Samuelson & Zeckhauser 1988·Kahneman, Knetsch & Thaler 1991·Kim & Kankanhalli, MISQ 2009
§4.2 · Mechanism: the body keeps the map
Users remember where, not what
Skilled interface use isn't thinking; it's procedural memory. Cohen and Squire showed it's a separate memory system from knowing facts, and Karni's fMRI work showed practice physically rewires motor cortex. After enough repetitions, reaching for a control bypasses conscious search entirely. That's why you can type without looking, and why Fitts blamed “human error” on design fifty years before we shipped software.
The spatial-memory literature sharpens it: Scarr, Cockburn, and Gutwin showed stable layouts let experts retrieve commands from remembered location, skipping visual search; the entire novice-to-expert transition depends on positions holding still. Labels, colors, and content are cheap to change. Positions are expensive.
Moving a button a user hits fifty times a day isn't a layout decision. It's a neurological misfire.
EvidenceCohen & Squire, Science 1980·Karni et al., Nature 1995·Fitts 1954·Scarr, Cockburn & Gutwin 2013
§5 · The turn
Don't generate interfaces. Generate within interfaces.
The evidence doesn't say never adapt. It says: change what the interface shows. Don't change where the interface lives. Watch the same generative power spent two ways:
§5.1 · Requirement 01
Structural consistency
Define a spatial grammar of fixed zones (nav, content, actions, status) that the model cannot re-decide. Generative freedom lives inside the zones. The header stays a header. The primary action stays where the hand learned it. This is Nielsen's “stable skeleton”; it's also how every platform that actually shipped in 2026 chose to work.
EvidenceNielsen 2000, Jakob's Law·Gajos et al., CHI 2008·Nielsen, July 2026
§5.2 · Requirement 02
Progressive disclosure over radical adaptation
Stage the rollout. New elements arrive visibly differentiated; removed ones collapse rather than vanish; the user can always trace the delta from last time. Findlater's group showed the mechanism works: gradual-onset “ephemeral” menus beat static ones when predictions were good, without the disorientation of reordering. And when YouTube replaced its entire design in 2017, it shipped the new look as an opt-in preview with revert for months before making it default. The biggest sites already treat radical change as a staged negotiation. Your model should too, remembering from §3.3 that “gradual” means signaled, not snuck.
EvidenceFindlater, Moffatt, McGrenere & Dawson, CHI 2009·YouTube 2017 rollout·Brock & Quigley et al., CHI 2018
§5.3 · Requirement 03
Accuracy over novelty
Gajos made accuracy the strongest lever we have: getting the right controls in front of the user beats getting interesting ones. The 2026 evidence adds a twist: Romero's team found AI-generated interfaces score well on usability and poorly on originality. Models generate conventional layouts. So the real predictability threat isn't wild design. It's variance: a systematic review of generative no-code tools found identical prompts producing substantially different interfaces across runs of the same tool. A generated checkout still needs cart, shipping, payment, confirmation, in that order, every time it's yours.
The model doesn't need taste to satisfy this requirement. It needs a memory of what it gave you last time.
EvidenceGajos et al., CHI 2008·Romero et al. 2026·Computers 15(4), 2026·AlignUI 2026
§5.4 · Requirement 04
Preserve escape routes
Remember the hidden cost from §3.3: the more accurately an interface adapts to your current task, the less of the system you're aware of. An AI optimizing for your immediate intent is an AI quietly deciding what you'll never discover. The remedy is structural: a stable “show everything” escape hatch, in a fixed location, that restores the comprehensive view, plus an undo for any layout the model chose for you. The endowment effect says a learned interface is a possession; revert is how the user keeps the title.
EvidenceFindlater & McGrenere, CHI 2008·Findlater & McGrenere, IJHCS 2010·Kahneman, Knetsch & Thaler 1991
§5.5 · Requirement 05
Respect motor memory
EvidenceFitts 1954·Karni et al. 1995·Cockburn, Gutwin & Greenberg, CHI 2007·Scarr et al., CHI 2012
§5.6 · Requirement 06
Transparent adaptation
An unexplained adaptation feels like a colleague rearranging your desk over lunch. The fix costs one line: “Showing editing tools because you're cropping.” Festinger's cognitive-dissonance work explains why this matters more than it looks: when reality contradicts “I know how this works,” people resolve the conflict by rejecting the tool, not by relearning it. An annotation converts the change from a violation of the user's model into an update to it. This is also where my own conviction sits: a system earns trust by being predictable in how it behaves, and explicit about what it doesn't know.
EvidenceFestinger 1957·Fraunhofer IOSB, GenUIn·Altaffer, “Three convictions”
§5.7 · Requirement 07
Convention as constraint layer
Jakob's Law: users spend most of their time in other interfaces, so they arrive carrying a composite model of how software works. A generated interface doesn't get to opt out of that model. It inherits it. Nielsen put a number on when deviation is justified: only when the alternative measures at least 100% higher usability. Not ten percent better. Twice as good. Swipe-back, pull-to-refresh, how a toggle behaves, what a modal does: these are immutable rails. Encode them as constraints the generator physically cannot cross, the way OpenAI's and Anthropic's app frameworks already do with reviewed component vocabularies.
§5.8 · Requirement 08
Adaptation pacing
EvidenceLavie & Meyer, IJHCS 2010·Findlater & McGrenere, CHI 2004·Yerkes & Dodson 1908
§6 · Compliance: the constraint layer is already winning
Nobody ships freeform generation to a billion people
Look at what the platforms actually built, versus what the demos promised. OpenAI renders developer-authored, reviewed components; the model chooses which, never what. Anthropic's MCP Apps run designed UIs in sandboxes. Vercel's AI SDK moved from freeform streaming toward typed tool-to-component mappings. Thesys generates a constrained spec into a fixed design system. In the academy, Cao and Xia's CHI 2025 system generates a task data model first and derives the interface by rule: consistency by construction.
Every one of these is REQ‑01 and REQ‑07, implemented independently, by teams that mostly haven't read the adaptive-interface literature. They rediscovered it the expensive way. The remaining six requirements are where the next generation of these systems will be won. They're cheap to adopt now, and brutal to retrofit after your users have learned three different layouts.
Constraints aren't the caution against the promise. They're the mechanism that delivers it.
EvidenceOpenAI 2025·Anthropic 2026·Vercel AI SDK 5·Thesys C1·Cao & Xia, CHI 2025
§7 · Open issues: logged for future revisions
What this document doesn't settle yet
| Issue | Question | Status |
|---|---|---|
| OI‑01 | Cross-session memory. How much should the system remember between sessions? Full continuity calcifies early adaptations; full reset discards everything the user taught it. | Open |
| OI‑02 | Multi-device coherence. How much consistency does a user need between the phone, tablet, and desktop versions of a generated interface? | Open |
| OI‑03 | Shared screens. Personalization that helps the owner confuses everyone watching (screen shares, pairing, support calls). Largely unstudied. | Open |
| OI‑04 | Accessibility. Ability-based generation (SUPPLE++) made motor-impaired users measurably faster: genuinely promising, and it demands capability models most teams don't have. | Open |
| OI‑05 | Measurement. Usability metrics assume a stable interface to measure. Validated frameworks for evaluating a moving one don't exist yet. | Open |
EvidenceGajos, Wobbrock & Weld, UIST 2007 (SUPPLE++)·Lee et al., DIS 2025·Ink & Switch 2025·Okopnyi et al. 2024, “Against Generative UI”
Annex A · References: every claim in this document links to one of these
One more thing, about the document you just read. The title block never moved. The clause numbers told you where you were. Every sheet changed completely (evidence tables, live figures, normative fields) inside a structure that never did. You read twenty-two screens of dense material and never once wondered where you were.
That's the whole argument.
Empirical HCI
- Gajos, Everitt, Tan, Czerwinski & Weld — Predictability and Accuracy in Adaptive User Interfaces, CHI 2008
- Lavie & Meyer — Benefits and costs of adaptive user interfaces, IJHCS 2010
- Findlater & McGrenere — Static, adaptive, and adaptable menus, CHI 2004
- Findlater & McGrenere — Screen size, awareness & adaptive GUIs, CHI 2008
- Findlater, Moffatt, McGrenere & Dawson — Ephemeral adaptation, CHI 2009
- Sears & Shneiderman — Split menus, ToCHI 1994
- Cockburn, Gutwin & Greenberg — A predictive model of menu performance, CHI 2007
- Scarr, Cockburn & Gutwin — Spatial memory in user interfaces, FnT HCI 2013
- Brock, Quigley et al. — Change blindness in proximity-aware mobile interfaces, CHI 2018
- Gajos, Wobbrock & Weld — SUPPLE++, ability-based UI generation, UIST 2007
- Scarr, Cockburn, Gutwin & Bunt — Improving command selection with CommandMaps, CHI 2012
- Findlater & McGrenere — Beyond performance: feature awareness, IJHCS 2010
Cognitive science
- Fitts — Information capacity of the human motor system, 1954
- Hick — On the rate of gain of information, 1952
- Miller — The magical number seven, 1956
- Sweller — Cognitive load during problem solving, 1988
- Kahneman & Tversky — Prospect theory, 1979
- Samuelson & Zeckhauser — Status quo bias, 1988
- Kahneman, Knetsch & Thaler — The endowment effect, 1991
- Cohen & Squire — Knowing how and knowing that, 1980
- Karni et al. — Motor cortex plasticity during skill learning, 1995
- Simons & Chabris — Gorillas in our midst, 1999
- Yerkes & Dodson — Strength of stimulus and habit-formation, 1908
- Festinger — A Theory of Cognitive Dissonance, 1957
- Kim & Kankanhalli — User resistance to IS change, MISQ 2009
Generative UI, 2024–2026
- Leviathan, Valevski, Natchu & Matias — Generative UI, Google Research 2025
- Cao & Xia — Generative and malleable UIs from task-driven data models, CHI 2025
- Chen et al. — Generative interfaces for language models, 2025
- Lee et al. — Towards a working definition of generative UI, DIS 2025
- Romero et al. — Usable but conventional, 2026
- Liu, Sra & Xiao — AlignUI, 2026
- Okopnyi, Nordberg & Guribye — Against Generative UI, 2024
- Litt, Horowitz, van Hardenberg & Matthews — Malleable software, Ink & Switch 2025
- Design behaviour and interface consistency in generative no-code tools, 2026
Industry & discourse
- Moran & Gibbons — Generative UI and outcome-oriented design, NN/g 2024
- Nielsen — End of web design (Jakob's Law), 2000
- Nielsen — The need for web design standards, 2004
- Nielsen — Generative UI from Gemini 3 Pro, Nov 2025
- Nielsen — Mid-year reality check, July 2026
- OpenAI — Apps in ChatGPT, 2025
- Anthropic — Interactive tools in Claude (MCP Apps), 2026
- Vercel — AI SDK 5, 2025
- Thesys — Generative UI architecture
- YouTube — A sneak peek at YouTube's new look and feel, 2017
- Fraunhofer IOSB — Generated user interfaces (GenUIn)
- CHI 2026 workshop — What does generative UI mean for HCI practice?
Lawrence Altaffer
Product design leadership · Richmond, VA