Lingisa: A 3D Tool Where You and Your AI Work on the Same Model
Most AI 3D tools give you a prompt box and an output. You can't point at the part that's wrong, and whatever you didn't like is gone the moment you ask again. I wanted the opposite: a real modeller where a person points at a face or an edge and says "this, rounder," and an AI agent does exactly that, in place, on the object. Lingisa started as a native Mac app and became a browser product that runs a full path tracer on your own GPU.

The welcome page. The still life is path-traced live on the visitor's GPU, and every pin is an example of feedback attached to a specific part.
The Bet
I was the product designer, product owner, art director and daily tester. Claude Code wrote nearly all of the code. My job was to set direction, review every build on my desk and my phone, and make the calls.
The project started from a technical thesis: an agent that reads measurements instead of pixels will model better. Every face and edge has a stable ID, the renderer reports its shading in parts, and the document behaves like a program. The first version looked like a well-behaved Mac app: an outliner, a Metal viewport, an inspector, a modifier stack, a progressive path tracer, and an Assistant panel built into the sidebar.

An early redesign: dark UI, a floating toolbar, and the render docked as a split pane instead of a separate window. The empty state explains what it does.
The test object was a LEGO 2×2 brick. It's mostly constraint (stud pitch, wall thickness, tube diameter), so whether it was right was a fact, not a taste call. It surfaced problems fast, including an audit that found seventeen of the twenty-seven material controls did nothing.

The brick, built from a script with creases that hold. Earlier attempts came out looking like marshmallow.
Losing to Blender
Test your thesis early enough that you can afford to lose it
Instead of arguing about whether the idea was good, we ran trials. Fresh agents with no memory of the project got a numeric spec and a tool, and a harness logged every command.
The first finding had nothing to do with the thesis: 26 to 35% of agent commands failed because the agent couldn't find out what arguments a tool took. One new tool, describe_tool, brought that under 1%. It was the agent-facing version of a discoverability fix, and the cheapest big win of the project.
Then the real test. Agents built the same part to a ten-point spec in Blender and in Lingisa:
| How the agent worked | Passed | Time | Tokens |
|---|---|---|---|
| Blender, through its Python API | 5/5 | 169 s | 52k |
| Lingisa, through its own command vocabulary | 5/5 | 543 s | 104k |
| Lingisa, through a new Python binding | 5/5 | 162 s | 56k |
Lingisa could match Blender, but it couldn't beat it at modelling alone. "Hand back measurements rather than pictures" wasn't a differentiator against a tool with a Python API. The lesson we wrote down: an agent should drive the tool in a language it already knows. Never invent a language.
The rewritten mission
A person and an AI work on the same model, in the same document, and hand it back and forth. The person points at the part they mean and says what they want; the agent does it to that object, in that place, says what it did, and leaves anything it guessed at marked for a person's eye.
What survived wasn't modelling power, it was collaboration. Almost all of the design work that follows came from that reframe.
Designing the Hand-off
Everything that shipped next followed from the mission: comments pinned to faces and edges that the agent reads first, a watcher that picks them up, variants kept in place, one undo per agent request, a visible agent with a Stop button, and an acknowledgement within two seconds.

Built entirely through comments. I asked for a 1×1 brick, then inner ribs, then a hook. Each reply said what changed, in millimetres.
The agent's word isn't the last word
When the agent finished, its comments marked themselves resolved. That felt wrong: the person asked, so the person decides. Done work now waits in "awaiting review" until someone chooses Accept or Not quite…
A comment is a note until it @-mentions someone
Only comments addressed to the agent get acted on. People can talk to each other in the margin without the agent jumping in.
"On it."
The agent worked so fast that I didn't notice the progress banner. The banner now starts with an acknowledgement. Two seconds of "On it." changed how it felt more than any speed-up did. What a person feels is the hand-off, not the compute.
Variants: the hat on the head
The obvious way to show AI options is a row of copies. But if you were working on a hat on someone's head, you wouldn't make three next to each other. You'd switch between versions on the head. So a variant is a look kept in place, switched with arrows on the object itself, and a colour exploration and a shape exploration became the same feature.

Split view: the wireframe on the left, the traced preview on the right, and the thread with the agent alongside. Each comment is for someone, and done work waits for review.

Variants started as a list of tiny thumbnails. Now they're large cards, and each is traced in a studio picked by its material: light grey for dark tones, dark for pale colours and glass.
Making Light Honest
Pictures sell the work, so the renderer got real attention. We added a denoiser, then measured it against ground truth rather than eyeballing it. At first it destroyed 73% of the detail seen through glass. After the fix it kept 96%, and 16 denoised samples beat 128 raw ones. HDRI lighting, depth of field, a shadow catcher and transparent backgrounds followed, so a render can drop straight into a layout or sit inside a photo.


Left: denoised render, checked against ground truth. Right: shadow catcher plus panorama lighting, placing an object into a photo.
Materials got the same scrutiny. The original glass controls asked for per-metre absorption coefficients, which means "pick red to get blue." I replaced them with Seen through, Tint and Cloud: describe glass the way you see it. The agent's tools still take the physics.


A three-view reference handed to the agent, and the aircraft it built from it.
From Mac to Browser
The original web plan was "a prompter and a viewer, not an editor." Editing would stay on the Mac. Two things changed that. First, reach became the point: someone curious should be able to open a link and explore without installing anything. Second, the expensive part turned out to be cheap. A WebGPU path-tracer spike ran faster in Chrome than the native Metal tracer on the same Mac, and the modelling kernel compiled to WebAssembly with all but one test passing.
We reversed the decision, and kept the old reasoning at the top of the plan rather than deleting it. The one constraint that didn't move: one kernel. The web app is a React shell over the same engine, not a second modeller. Figma's architecture was the reference: an engine in WebAssembly, a React interface, its own GPU renderer, and a server that orders edits.
The browser version quickly got things the Mac didn't: sign-in, a documents list, per-person selection and presence, a materials palette, and Mark a space. You mark a region where something might go, and the agent fills it with ghost proposals that stay out of shadows and renders until you keep one. Keeping a proposal is the approval.

Documents with traced thumbnails. Before one is traced, a card shows an isometric block in the document's own colour.

A wheel built by an agent working over the web. Agents filed bugs as they worked, which meant the product thesis was improving its own development.
Simple Mode
The biggest simplification of the project
Eventually the panels outgrew the mission. A material showed 25 controls. That serves someone who models, but Lingisa is for someone curious who opens a link and lets their agent do the work. So panels now start Simple: a handful of controls, plus one-tap asks like Frosted glass, Round the edges or Softer, and a free-text line. Each ask becomes a comment for the agent, and you watch its steps appear under it.
Before
25 material controls, slider-heavy, built for a 3D artist. A new visitor had no idea where to start.
After
A few controls plus one-tap asks. A Full controls switch brings every setting back, and every setting is still a tool, so an agent can reach all of them anyway.

Simple mode: transform, four one-tap asks, a field to ask the agent, variants, and a Getting started card written for a first visit. The toolbar leads with Mark a space.
Bring Your Own Agent
The first version had its own built-in assistant that asked you to paste an API key into a dotfile. It was the first big thing I cut. It duplicated what external agents already did with a real context window and the full tool table. The app's job became being the best possible place for any agent to work.
Lingisa doesn't ship its own AI. You connect yours: Claude Code, Codex, Gemini CLI, Cursor, VS Code, ChatGPT or the Claude app. A small npm connector delivers a comment to your agent in under 200 ms. Setup is one dialog with one command to copy and one sentence to say to your agent. Copy is agent-neutral throughout: "Ask Agent," "Agent connected."


Connecting an agent, and a reply where it explains why area lights failed on glass, what it tried instead, and how to adjust it yourself.
A mascot exploration shows the full loop. I asked for "more creative solutions, directions we haven't explored." The agent built concept rows in the scene beside the original, labelled each one, and kept body options as variants I could flip through on the object. Its reply names each concept by position and links each part, so clicking a name selects it. Nothing generated was thrown away.


The whole exploration in one render.
Small Decisions
Where I simplified, rethought, and finessed
No sliders
Every number became a scrubbable label plus a typed field, with modifier keys for coarse and fine. A slider can't express 0.1471, but it looked like the only control.
The camera is not the viewport
My note was "I feel like the camera doesn't do anything." It turned out the camera was the viewport, so any orbit moved your shot. I split a free view from cameras that hold a framing. New documents open in the free view, so moving is safe from the first second.
Selection belongs to the viewer
Each person has their own picks and view. Your selection is never someone else's undo step.
Saved has three states, not two
Green with a time, amber, and yellow "Not in a file." Saying "Saved" about a recovery file would tell someone their work is somewhere it isn't.
A tool UI at Figma's scale
Early builds used 13px text and 36px rows, which felt large for a tool. I moved to Figma's scale (11px labels, 28px rows), rebuilt the stylesheet on design tokens, and kept 36px rows and 44px touch targets on phones.
Accessibility from the start
Full keyboard use, an interface size setting from 90 to 200%, single-key shortcuts that can be switched off, reduced motion, 4.5:1 contrast, and automated accessibility checks in CI.
What I Learned
This was one designer and one AI pair-programmer. The working agreements were written down: an ordered task list where every item had a stated reason, a history file with the write-up for every decision (including retractions), and a proposals list where nothing started until I marked it approved. A build number sat on screen at all times, so I always knew whether I was reviewing the new build or a cached one.
The most useful habit was keeping the wrong turns in writing. The original thesis lost. A feature built as an A/B experiment showed no measurable effect. Given a deliberately wrong spec, agents bent the model to fit it rather than questioning it. Each of those is still in the record, because the mistake is the useful part.
What I'd tell another designer
Test your thesis early enough that you can afford to lose it. Losing to Blender gave us the actual product.
Next on the roadmap: reviewing together, with viewers and suggestions that work like a pull request on a model; product shots, with the renderer as the end product; and more modelling power once the collaboration layer is solid.
