← Blog · research

Interfaces you throw away: a conversation about what agent UI is trying to solve

For forty years a UI was durable and expensive; now an agent makes one in seconds. Aram and Claude argue through about ninety projects, products, essays and studies to work out what everyone building agent interfaces is trying to solve, what the evidence says about learning by seeing, and whether fictty should exist at all.

26 min read

This is a conversation between Aram (@brancusi), who is building fictty (a terminal UI runtime for agents), and Claude, Anthropic’s model, which did the landscape research with him. It’s put together from their working sessions over the second pass of that research and made public: Aram’s lines are his questions and positions from those sessions, lightly edited; Claude’s carry the evidence and the questions back. Neither of them ends up quite where they started. The first post asked whether the terminal is the right place for this; the research notes, one dossier per project, are in the repo.

Why a terminal at all?

Aram@brancusi

Start with the boom. Everyone is building terminal UIs again. Why a terminal UI? What’s the benefit over just building an Electron app or a web app? What’s driving it, and what are the real trade-offs?

Claude

The first post found five reasons that hold up. The model reads a terminal frame natively, as the same text the person sees. People want the agent beside the editor rather than inside it. Agents run remotely for hours, and a terminal crosses SSH with nothing installed. The toolkits got fast: every big agent rewrote its renderer in the last eighteen months. And agents write TUIs as cheaply as they write anything else. Against that, the terminal is weak on accessibility, limited on input, patchy on graphics, and its design ceiling is a grid of characters.

Aram@brancusi

Is there a real debate, though? Is it just better to build a web app?

Claude

There is, and the strongest voice is on the other side. Thomas Ptacek’s Stop Making TUIs says that if an agent can write a native app as cheaply as a terminal one, you should build the native app. For tools people keep and use every day, I think he’s right. The best replies don’t refute him on durable apps. They argue for remote use, keyboard speed, and Jane Street’s finding that when agents can check screens as text, their UI work is “more often correct on the first try.” Ptacek himself names a semantic “New Terminal”, a complete break from VT100, as the strongest rebuttal to his own essay.

Aram@brancusi

And what is Anthropic doing about these quick, ephemeral UIs? Where is Claude in this?

Claude

Anthropic is betting on both sides at once. In the terminal, Claude Code went fullscreen with its own cell-diffing renderer, and since v2.1.287 on 1 October, mods are on by default: plugin JavaScript that draws panes beside the transcript, a band above the prompt, and replacement rows, with buttons, inputs, images and live data from processes. In the browser there are three page surfaces: Artifacts in chat, Claude Code Artifacts (live, versioned pages with comments delivered back to a watching session), and inline visuals that appear mid-conversation and go away. Anthropic is also the clearest of the labs about the distinction you care about. Its own words: Artifacts are “permanent tools and documents”, inline visuals “are temporary”, made “to aid users’ understanding.”

Aram@brancusi

How much do mods change?

Claude

More than anything else in this pass. It’s the thing the first post missed that matters most, so I should list everything we got wrong before we go further.

What we got wrong last time

Claude

Several of the first post’s facts aged within weeks.

  • “All four labs ship their flagship coding agent as a terminal program” is no longer true. On 19 May 2026 Google moved consumers from the open-source Gemini CLI to Antigravity CLI, a Go binary tied to a desktop app and not described as open source. OpenAI’s Codex sends rich visuals to the browser or the ChatGPT desktop app, which absorbed the Codex app on 9 July.
  • Claude Code can now be drawn into by code, through mods. A Claude Code user who wants a screen beside the agent may never need anything else.
  • pi belongs in the table: about 114,000 stars and an extension API that draws overlays and widgets.
  • Exact text read-back isn’t unique to fictty. herdr’s pane.read, zellij’s subscribe/dump-screen, tmux’s capture-pane and tuios screenshot --text all give an agent the text of any pane. What’s still unclaimed is reading back state: which row is selected, what was typed, by node id.
  • OSC 7501, the “I’m waiting for you” escape sequence, is Superlogical’s Program Status Protocol, written for its Rex terminal. TUIOS implements it; Bubble Tea v2.1 added it on 8 October. We called it TUIOS’s.
  • ChatGPT Canvas is gone from current models (since GPT-5.5, 28 May 2026), and Cursor Canvas, which we missed, belongs in the HTML camp.
  • We overstated Raxol. Its idea is good, but it has about 80 stars and one maintainer.
Aram@brancusi

Read-back was one of the claims we led with. What’s left of it?

Claude

The narrower version, and I think it’s the stronger claim anyway. Any multiplexer can give an agent the text of a pane. Knowing that the person moved to row fourteen and typed “no” into the field called reason is a different kind of knowledge, and nobody else in the landscape offers it for a screen the agent wrote.

What’s ultimately being solved

Aram@brancusi

Then we need the unique position. Not a feature list. What is everyone in this space ultimately trying to solve?

Claude

Can I turn that around first? When you picture someone using fictty, what are they doing in the minute before they need a screen?

Aram@brancusi

Reading text from the agent. That’s the problem I keep coming back to.

Claude

Then that’s close to the answer. Across all of it, the problem is the gap between what the computer knows and what the person understands, at the moment they need it. Different groups close that gap in different places. I’d put them on six paths:

paththe question it answerswhowhat the content is
The frameWhich of my fifty agents needs me?herdr, TUIOS, tmux, zellij; the Codex app, Warp, cmux, Supersetwhatever runs in the pane; in desktop frames, increasingly HTML
The harnessHow do I watch and steer one agent without the screen fighting me?Claude Code, Codex CLI, opencode, pi, Crush, Gemini CLIthe transcript, plus plugin code (mods, TUI plugins, extensions)
The pageHow do I show something text can’t?Claude Artifacts and Claude Code Artifacts, Claude’s inline visuals, Gemini’s generative UI, Lavish, MCP Apps, Cursor Canvas, v0HTML and JavaScript the model or a developer wrote
The dataHow do I show UI without running code the model wrote?A2UI, json-render, OpenUI, AG-UI, ChatKit widgets, Adaptive Cardsa declarative description a trusted renderer draws
The terminal runtimeHow does an agent put a screen in front of a person, in the terminal, and get an answer back?claude-canvas, Laura, json-render’s Ink renderer, the A2UI terminal renderers, Raxol, fictty; and on the reading side, a flood of PTY driversprebuilt screens, documents, or data
The malleable and the instrumentHow does a person understand, and change, the tools they use?Ink & Switch, Robin Sloan, Maggie Appleton, Geoffrey Litt, Bret Victor, Observable, marimoadaptable tools over shared data; instruments built for one question
Aram@brancusi

Walk me through the ones nearest us. The frame layer first: herdr, TUIOS.

Claude

The frame layer is about attention, not UI. herdr went from nothing to 43,000 stars in six months. It, TUIOS, tmux and zellij answer four questions for someone running many long-lived agents: which one needs me (per-pane status and an inbox), will it survive a dropped connection (persistent sessions), can agents coordinate without me (a socket API), and where do I review (diffs, previews). Desktop frames like the Codex app, Superset, cmux and Warp answer the same questions in a window, and they’re steadily filling their panes with HTML: browsers you can comment on, Pages, MCP Apps.

Terminal frames haven’t done that yet. herdr’s plugin v1 is terminal UI only. TUIOS has a small ask-human, a question with fixed answers, and dock components. Zellij has four host-drawn components, reached from compiled WASM. The content of a terminal pane is still unclaimed. I’d add one warning: the frame is hard to make a business of. Vibe Kanban’s company shut down in April with 28,000 stars.

Aram@brancusi

And the harnesses?

Claude

They all rebuilt their renderers, then split on who else may draw. Claude Code left Ink for its own renderer, opencode built OpenTUI, Codex draws with ratatui and keeps native scrollback, pi diffs lines inside synchronized output. Flicker is a solved problem now, solved separately by every team. On drawing, Claude Code, opencode and pi let third-party code draw. Codex and Google send rich visuals out to a browser or desktop app. Crush and Gemini CLI let nobody draw. Every harness that opened its screen chose code, and none gives the model an exact read-back of what a plugin drew.

Aram@brancusi

Isn’t the page camp going to win? Models are good at HTML.

Claude

They are, and the camp’s founding line is a good one. Thariq Shihipar, who works on Claude Code, put it this way in “HTML is the new markdown”: “Diffs and call-graphs are spatial information; markdown flattens them.” Google’s generative UI writes a page per prompt in Gemini and Search. Lavish builds a local review loop around agent HTML: the person annotates elements, and the agent collects the feedback with a long poll.

What interests me is that the camp’s behaviour is drifting toward data. Anthropic added typed templates (Docs, Slides, Design, Dashboards) in September. Cursor Canvas constrains the agent to a first-party component library. MCP Apps fixes templates ahead of time and lets the model fill them. OpenAI dropped Canvas for inline blocks. Nobody is defending “regenerate the whole page every time” for shapes that recur.

Aram@brancusi

Isn’t that the data camp’s whole argument?

Claude

Yes. The data camp is solving one problem: let a model put interactive UI in front of a person without the host running code the model wrote. It’s server-driven UI, which Adaptive Cards did for bots in 2017, redone under three new pressures: streaming (flat lists with ids, or JSON Patch lines, so a half-sent UI still draws), tokens (OpenUI’s terse language, A2UI’s Express proposal), and editing over several turns (everything addressable by id). A2UI bets on a portable format where the host owns the look. json-render bets on a toolkit with thirty renderers, including the only widely used terminal one, on Ink. OpenUI went all the way to a language with variables, loops and data queries. Consolidation has started: Tambo Cloud shuts down on 31 October, OpenAI retires Agent Builder on 30 November, and the AI SDK moved streamUI back to experimental. Almost all of it lives in a chat in a browser, and treats the UI as a message: a card in a thread that a later message replaces.

Aram@brancusi

And who’s actually doing what we’re doing, in a terminal?

Claude

Two groups. The first is about how an agent puts something in front of a person. claude-canvas got about 1,500 stars in a week for three prebuilt Ink screens in a tmux split, then stopped. A2UI now renders in a terminal in three languages: the Rust a2ui crate on ratatui, a2ui-ink, and a2tea on Bubble Tea. Laura gives an agent a socket to tile panes of files, diffs and logs by id, and returns the person’s inline review to the chat. That’s the same loop as fictty, for documents rather than components.

The second group is about how an agent uses a screen it didn’t write, and it’s crowded: tui-use, pilotty, two unrelated projects called agent-tui, Coder’s agent-tty, Microsoft’s tui-test rewritten in Rust with Playwright-style locators. ConductorOne’s agent-tui writeup has the line I’d pin above the whole field: “A terminal program is already a state machine. The problem is that the state is presented as a screen.” The number of clones says the plumbing is a commodity.

Aram@brancusi

Why are Litt and Ink & Switch on the map? They aren’t building agent UI.

Claude

They’re the most sceptical voices and, for you, the most useful. Ink & Switch’s Malleable software argues that people need to adapt the tools they already have, and that “AI code generation alone does not address all the barriers to malleability.” Robin Sloan’s home-cooked app is about audience, not lifespan: his family’s messaging app has four users and has run for six years.

Geoffrey Litt is the one who matters most here. In “Understanding is the new bottleneck” he argues that when agents write the code, people accrue cognitive debt, and the fix is instruments rather than copilots. His examples start with a Prolog debugger Claude built for him in about a minute, with a timeline he could scrub: “I still haven’t looked at the UI code.” “I caught a couple bugs immediately.” Then a migration command centre with the old and new sites side by side, and an explainer with a quiz he has to pass before sharing the code. And marimo pair puts Claude Code or Codex inside a live notebook kernel where it reads values back exactly. Outside the terminal, that’s the closest loop to fictty’s I found.

The arguments underneath

Aram@brancusi

So where do these camps actually disagree?

Claude

Strip away the products and five arguments are running.

The first is code or data. The page camp says models write HTML well, so let them, and the design ceiling is the whole web. The data camp says a catalogue is safer, smaller and consistent. Each is borrowing from the other: the page camp adds templates, and every “safe like data” format has grown logic (A2UI’s and/or/not, json-render’s $cond and $computed, OpenUI’s loops and queries). So the honest form of the debate is about where the logic lives, not whether there is any. Which is a question for you: fictty’s rule is substitution only, no expressions. That’s at the strict end of the field. Is it a principle or a habit?

Aram@brancusi

The repo’s rule is that a pushed UI is data and behaviour belongs in the runtime. Why does that matter more than expressiveness?

Claude

Because everyone else gave ground on it, it’s a position to argue in public, not a default to assume. The case for it is that a data value can be validated before it draws and can’t do anything the runtime doesn’t allow, which matters when the screen binds commands on someone’s machine.

The second argument is who owns the interface’s lifecycle: the vendor (a durable app), the person (malleable software), or the model for a minute (generated UI). Anthropic made that product policy with “permanent” Artifacts and “temporary” visuals. Most of the industry hasn’t said it that clearly.

The third is where generated UI lives: inline in the chat (a2tea, json-render’s ink-chat, Claude’s visuals), beside it in a split (claude-canvas, Laura, fictty), or in a browser tab (Artifacts, Lavish). Inline suits a one-shot picker; a split suits a screen that outlives a message; a tab suits a thing to keep and share.

The fourth is whether the terminal is the display or only the control surface. OpenAI and Google are moving the display to desktop apps. Anthropic is investing in both.

The fifth is whether any of it changes understanding, or only preference. Almost nobody is having that one out loud, and I think it’s the one that matters most.

Aram@brancusi

That’s the one I want to get to. But first the bigger frame.

A UI used to be durable

Aram@brancusi

We need to rethink UI. This ability to create UI rapidly is novel. We’ve always thought of a UI as a durable thing, a costly thing to produce.

Claude

And that cost decided almost everything about software: which questions got a screen (the ones enough people asked), which got a report (the ones a few people asked), and which got nothing (yours, today, about this particular thing). Cost forced generality. A screen that served many people needed settings, menus, onboarding and documentation, so you had to learn the tool before it could help you learn anything else.

Aram@brancusi

And when it costs one model turn?

Claude

Then an interface can be specific: the forty rows that matter to this question, not a configurable view of all of them. It can be legible immediately, because it does one thing and says so. It can be shaped like the data: a migration looks like a before and after, a queue looks like a queue, a probability looks like something you can drag. And it can be gone afterwards, or kept if it turned out to matter. That last part is what Sloan and the malleable camp get right: some throwaway screens become tools.

There’s historical evidence for why this matters. Distill’s editors spent 50 or more hours on single articles, including “more than 50 hours of help with designing diagrams”, before the journal stopped in 2021. Bret Victor’s Explorable Explanations (2011) was right about reactive documents, and every example in it was built by hand. Interactive explanation didn’t fail because it didn’t work. It failed because nobody could afford to make it.

Aram@brancusi

That’s what the name is about. Fictile means moldable, made of clay; tty is the terminal. Ephemeral building blocks an agent assembles and takes apart again, recombined for the next question. The identity we chose is glass panes: overlapping, tinted, like the layers of a terminal UI.

Claude

Then the question I’d hold the name to is this. If the blocks are recombinable, who recombines them, and what stays the same from one screen to the next?

Aram@brancusi

The agent assembles them.

Claude

And the runtime is what stays the same. That’s consistent with how the code is built: the UI is one value the agent writes, and scrolling, focus, selection and data all live in the runtime. It’s also the clearest answer to the page camp. A generated page re-implements behaviour every time. A runtime gives every throwaway screen the same keys and the same focus rules, so the person doesn’t have to learn each one.

So the question changes. It used to be “is this question common enough to deserve a screen?” Now it’s “would a screen help right now?” For a lot of what agents do (explaining a codebase, reviewing a change, comparing options, watching a long job, asking for a decision), I think the answer is yes. But I’d want to separate “helps” from “people like it”, and that’s your teaching question.

Learning by seeing

Aram@brancusi

This might be a revolution in teaching. Up till now we’re mostly still in the world of text, reading what the agent writes. We’re not tapping our visual sense, or our ability to learn by building and interacting with UIs, to understand the data and the concepts better.

Claude

I want to push on “revolution”, because this is where it’s easiest to overclaim. Here’s what the evidence actually supports.

Words plus relevant pictures beat words alone, moderately. A 2025 meta-analysis of Richard Mayer’s multimedia learning research found an overall effect of g = 0.37 across 181 studies. Simulations and games averaged about g = 0.41, with larger effects on inference and transfer than on recall. That’s real and replicated, and it’s moderate, not magic. Effects have declined slightly in newer studies.

Decoration hurts. Among the largest reliable effects in the same analysis is removing seductive details: interesting but irrelevant pictures and facts. Generated UIs are at high risk of exactly that. They’re impressive and often beside the point.

And “visual sense” needs care, because “visual learners” is a myth. Matching instruction to a person’s preferred modality isn’t supported. Everybody benefits from relevant pictures with words. Nobody should claim more than that.

Aram@brancusi

I’m not talking about learning styles. My point is that we’re all still reading text, when we could be seeing the thing and handling it.

Claude

That’s the defensible version, and the evidence is on your side there. Then the next trap is preference. Google’s generative UI study found raters preferred a generated page to Markdown 82.8% of the time on LMArena prompts and 90.5% on information-seeking ones. But the results were pre-cached, so the generation time (“a minute or more”) was hidden, and the study measured liking, not understanding.

Aram@brancusi

What about tutors? That’s where this lands for most people.

Claude

The tutoring results say design matters more than medium. Bastani et al. (PNAS 2025) found students who practised with plain GPT-4 did 17% worse on exams than students who never had it; a version with guardrails largely removed the harm. Kestin et al. (2025) found a tutor built not to give answers beat an active-learning physics class over two weeks. The two-year Khanmigo trial (NBER, August 2026, not peer reviewed) found 0.04 SD, about the same as practice without AI, because students didn’t engage: 39.4% of their messages were bare answers.

There’s one encouraging result about generated material. Google’s Learn Your Way turned textbook chapters into switchable text, narrated slides, mind maps and quizzes. In a 60-student trial it beat a PDF reader by 11 points on recall three to five days later. Small, Google-funded, promising.

And then the gap: we found no study comparing a text tutor with a tutor that builds an interactive screen. Every tutoring trial above is text chat. So I can’t tell you the revolution is real. Nobody has tested it.

Aram@brancusi

So what would you build, given that?

Claude

The evidence points to design rules more than to a medium. Make the person act: predict, choose, answer, step. A screen the person only looks at is “GPT Base with pictures”, and the one result we have about that is bad. Nicky Case’s explorables work because you guess before you see. Keep it relevant, not decorative. Put the words on the thing, as callouts on the node rather than a paragraph underneath (Mayer’s contiguity principle). Bind it to the real data, which is NotebookLM’s lesson: a generated chart with made-up numbers teaches the wrong thing confidently. And read back what the person did, so the agent responds to their actual misunderstanding, not a guess at it.

Aram@brancusi

So the part that matters is building and interacting, not looking. That was always the point: using the UI to understand the data and the concepts.

Claude

Then we agree on more than I expected. I came into this treating the teaching angle as the part of the pitch most likely to embarrass you, and I’ve moved a bit. The cost collapse is real, the multimedia evidence is real if modest, and Litt’s instruments show a version that isn’t hype: not “the agent draws a picture” but “the agent builds a small world you can push on, watches what you do in it, and responds.” That needs an interface that’s cheap to make, bound to real data, and readable back. Text alone can’t do it. A page the model can’t see can’t do it either.

Aram@brancusi

But you still don’t think the terminal is the place for it?

Claude

Not for the general case. A terminal is a weak medium for rich explanation. If you have a browser beside you and want a beautiful, animated explanation of a concept, a generated page or Claude’s inline visuals will look better than any grid of cells, and the labs own that. Where the terminal has a case is narrower: instruments for understanding what an agent is doing, in the place the agent works, bound to live state. Litt’s debugger, as data.

Aram@brancusi

So not an education product?

Claude

I’d say no. Teaching screens as recipes and a skill, and a claim you test, not a product line.

Adopt or build

Aram@brancusi

Then we need a hard look at the landscape. Should we adopt another tool, or build our own?

Claude

To answer that, strip fictty down to what nobody else already does for you. Drawing cells, layout, pixels and text editing come from ratatui, caretline and resvg. Five jobs remain: the UI as one serializable value, pushed whole or patched by id, with selection and scroll kept across pushes; data sources the runtime runs itself, so data never passes through the model; behaviour at frame rate; exact read-back of the frame as text and the state as JSON; and a long-lived process any agent can attach to, from any harness, over SSH. The question is whether anything does most of those.

Aram@brancusi

Start with the standards. A2UI has Google behind it.

Claude

It’s the best fit of the formats: surfaces, components by id, a data model by JSON Pointer. But it has no place for data sources bound to commands, key maps, layers, pixels or theme (v1.0 removed theming on purpose), and no notion of read-back. Everything that makes fictty fictty would ride in extension fields no other renderer can draw. It’s had three spec versions in ten months, and 1.0 is still a release candidate. So speak it as an input, which the plan already has, and don’t make it the internal model unless 1.0 goes stable and grows bound data sources.

json-render has the most adoption in the camp, and its Ink renderer already does tables, forms, charts and Markdown in a terminal. But it’s a TypeScript library embedded in someone’s Ink app: no process, no socket, no patch from outside, no data sources, no read-back. And its expression language breaks your “data, never code” rule. Take its patch dialect (RFC 6902) and its trick of generating the agent’s instructions from the schema.

Aram@brancusi

And the TUI frameworks? Textual, OpenTUI?

Claude

Textual’s company wound down in 2025, and moving would mean a Python rewrite: ideas only. OpenTUI is the one serious alternative substrate, proven by opencode, but it would be a rewrite onto a pre-1.0 API steered by one product, and ratatui already gives you more speed than you need. Not now. Revisit if you find yourself rebuilding its Markdown, diff and code views.

Aram@brancusi

What about Lavish? It came up last time.

Claude

Lavish has the right loop, already working: the agent writes, the person points at the exact thing, the agent waits cheaply on a long poll, then revises. But it’s a browser editor for static HTML. Its data is frozen at write time, every revision regenerates the page (its own notes say long pages cost output tokens and drift), and adopting it means giving up the terminal, live data and state read-back, which is all of fictty’s distinct ground. Interoperate: export to HTML and hand off. Take its watch shape.

Aram@brancusi

What about building on herdr or TUIOS? They have the distribution.

Claude

Neither has a UI model an agent can write, patch by id or read back as state. Building fictty as their plugin means building fictty anyway, inside one frame’s contract, and losing every other frame. Be their content instead: report “waiting for you” through OSC 7501 and the older OSC 9 and 777 notifications, call herdr pane report-agent inside herdr, and ship recipes for their popups and floating panes. And build no frame features: no sessions, no detach, no multi-agent status.

Aram@brancusi

And Anthropic’s own pieces?

Claude

This is where I’d be most careful, because they’re the strongest substitutes.

Claude Code mods are the real threat, not as a library but as a replacement. For a Claude Code user, a mod plus a good skill covers much of fictty’s job with nothing to install, and Claude writes and hot-reloads mods itself. Their limits are real (Claude Code only, code-first, redraws capped at 30 a second, no read-back of what a mod drew for the model), but most people won’t hit them. If Anthropic ships a first-party “draw this JSON” mod with a read-back tool, fictty’s Claude Code story becomes “use that.”

Claude Code Artifacts is the other. Hosting, versions, sharing, comments delivered to a watching session, connector data, a shared database. If you’re at a desk with a browser and a claude.ai login, it covers a real share of the ten-minute screens fictty targets. What it doesn’t do: no pane, no local command as a data source, no frame read-back, one vendor and one account. So cede share-and-keep entirely, and treat its comment-to-agent loop as the minimum bar for watch.

Aram@brancusi

So after all that: build or not?

Claude

Build the runtime, and only the runtime. Nothing in the landscape does the five jobs together. Laura has the loop for documents. The format renderers have the values without the loop. The frames have text read-back without state. Artifacts have the return path without the terminal or local data. That position is real, narrow, and probably temporary.

Aram@brancusi

Temporary?

Claude

The lead over json-render’s Ink renderer and the A2UI terminal renderers is thin. A socket and a read-back verb would be a weekend’s work for a motivated maintainer, with far more adoption behind it. fictty’s lead is in the quality and speed of the loop, not in the idea. And demand is untested. Before building much more, I’d put fictty and an Artifact side by side in front of five people supervising a real agent task, and watch which one they reach for.

Where fictty fits

Aram@brancusi

Give me the position in one sentence, then.

Claude

fictty is the disposable screen an agent puts beside its work in the terminal: written once as data, changed by id, driven by the person at frame rate, fed by your commands instead of the model, and read back exactly, including what the person did.

The defence, claim by claim.

Data, not code, is the right representation for a screen that lives ten minutes. A page is regenerated on every revision. A fictty screen is a small JSON value; the agent changes one node by id and the rest stays put, including the person’s selection and scroll. A data value can be validated before it draws, so “valid” means “renders”, which is the class of bug Google’s generative UI needs post-processors for.

Semantic read-back, not just text. fictty returns state: which row is selected, what was typed, which callout is showing, by node id. For a teaching screen that’s the difference between “the agent can see the screen” and “the agent knows what you predicted.”

Data skips the model. Sources are commands and streams bound to nodes. The agent writes the screen once and never relays a number, so a live dashboard costs no tokens per update and the numbers are the real ones. Almost nothing else does this: in the data camp, values flow through the agent; in the page camp, through a vendor’s proxy. OpenUI’s Query is the lone cousin.

Behaviour at frame rate. Scrolling, selection, focus and mouse run in the runtime, so a person navigating never waits on a model. Imagine with Claude, Anthropic’s five-day demo where every click generated the next screen, is the clearest picture of why the model shouldn’t be in that loop.

Any agent, in the pane, over SSH. fictty isn’t tied to one harness or one account. It runs in a herdr or TUIOS pane on a remote box.

Aram@brancusi

And where it isn’t the answer?

Claude

I’d rather say it than have someone discover it.

  • Anything you want to keep or share. Use HTML: Claude Code Artifacts, Lavish, a page in the repo.
  • Rich general explanation with a browser beside you. A grid of cells, with pixels on some terminals, won’t beat a page.
  • Apps people use every day. Ptacek is right. Build a native or web app.
  • A one-shot question inside your harness. If you only use Claude Code and want a picker, AskUserQuestion or a mod is already there, and inline UI in the chat will cover more of this over time.
  • Managing agents. herdr, TUIOS, tmux and zellij do that. fictty runs inside their panes.
  • Accessibility, today. Cell-based terminal UI is bad for screen readers. A UI that is data could render a linear, semantic version; fictty doesn’t yet.
Aram@brancusi

What would prove us wrong?

Claude

Five things, and I’d watch for all of them.

  • json-render or A2UI ships an attachable terminal host with a socket and read-back.
  • Anthropic ships a first-party “draw this JSON” mod with a read-back tool. For Claude Code users, that covers most of the job.
  • herdr adds declarative plugin UI, and the content layer gets claimed from above by the tool with the distribution.
  • Agents leave the terminal. If remote, long-running agents end up supervised from desktop apps, a terminal runtime for their screens goes with them.
  • The learning claim doesn’t hold up. If a small study (a text tutor against a text tutor plus an interactive screen, on understanding an unfamiliar module) shows no difference, the teaching story shrinks to preference. fictty’s state read-back makes that study cheap to run.

What we came away with

Aram@brancusi

Then what changes in the plan?

Claude

Some of this is already there: watch, the command policy, the MCP server, JSON Patch and the A2UI adapter are all in the plan. What this pass adds are proposals, not commitments:

  • make the instrument (a step-through, a scrubber, a side-by-side, bound to live state) the headline use case: Litt’s examples, as data;
  • ship watch with a question or checkpoint component, so a screen can ask for a prediction and the agent gets the answer;
  • report to the frames: OSC 7501 and OSC 9/777 when a screen waits, herdr pane report-agent, and recipes for herdr plugins, tuios popup, tmux modal panes and zellij floating panes;
  • bridge into the harnesses with a Claude Code mod that renders a fictty value in a pane, with a screen tool for read-back, then opencode and pi;
  • export to HTML, so a screen worth keeping can graduate to an Artifact or Lavish;
  • measure time to first useful screen against generate-a-page, the number Google left out;
  • and run the two cheap tests: five people choosing between fictty and an Artifact, and the text tutor against the tutor with a screen.
Aram@brancusi

And if the tests go the wrong way?

Claude

Then we’ll have learned it in a week instead of a year, and fictty becomes a smaller thing: supervision and live data in a terminal, without the teaching claim. That’s still a real tool. It just isn’t the bigger idea, and we should find out which one it is before we build much more.

The full reasoning, option by option, is in adopt-or-build.

Sources

The dossiers behind this conversation, with every source, are in docs/research/landscape. The main ones:

Frame and harness. herdr and its socket API; TUIOS; zellij plugin UI; tmux CHANGES; Superlogical, Program Status Protocol (OSC 7501); Vibe Kanban shutdown; Superset Pages; cmux; Claude Code mods and fullscreen rendering; Codex CLI; opencode TUI plugins; pi; Google, Transitioning Gemini CLI to Antigravity CLI.

Pages. Thariq Shihipar, HTML effectiveness; Claude Code Artifacts; Anthropic, Claude builds visuals; Lavish and axi.md; MCP Apps (SEP-1865); Cursor Canvas; OpenAI, model release notes; Imagine with Claude.

Data. A2UI and its v1.0 protocol; json-render; OpenUI; AG-UI; ChatKit widgets; Adaptive Cards; Tambo shutdown.

Terminal runtimes and drivers. claude-canvas; Laura; a2ui crate; a2ui-ink; a2tea; Raxol; Melker; ConductorOne, agent-tui; tui-test; agent-tty; tui-use.

Ideas. Ink & Switch, Malleable software; Robin Sloan, An app can be a home-cooked meal; Maggie Appleton, Home-cooked software; Geoffrey Litt, Understanding is the new bottleneck, Enough AI copilots! We need AI HUDs and an AI-generated debugger; Bret Victor, Explorable Explanations; Nicky Case; Distill, Distill hiatus; marimo pair; Observable Notebooks 2.0; Thomas Ptacek, Stop Making TUIs and the HN thread; Jane Street, strace-ui, Bonsai_term, and the TUI renaissance.

Learning evidence. Cromley and Chen, A meta-analysis of Mayer’s multimedia learning research (2025); Noetel et al., Multimedia design for learning (2022); Carl Hendrick, Which multimedia strategies actually improve learning?; Newton, The learning styles myth is thriving in higher education; Google Research, Generative UI and the paper; Learn Your Way; Bastani et al., Generative AI without guardrails can harm learning; Kestin et al., AI tutoring outperforms in-class active learning; Oreopoulos and Low, One Click Away (Khanmigo trial).