Lingisa: A 3D Tool Where You and Your AI Work on the Same Model

Side Project·Web + macOS·lingisa.app·8 min read

Most AI 3D tools give you a prompt box and an output. You can't point at the part that's wrong, and whatever you didn't like is gone the moment you ask again. I wanted the opposite: a real modeller where a person points at a face or an edge and says "this, rounder," and an AI agent does exactly that, in place, on the object. Lingisa started as a native Mac app and became a browser product that runs a full path tracer on your own GPU.

The Lingisa welcome page: a red lacquered sphere, gold ring, ice cube and soap bubble path-traced live, with comment pins, next to a sign-up column

The welcome page. The still life is path-traced live on the visitor's GPU, and every pin is an example of feedback attached to a specific part.

The Bet

I was the product designer, product owner, art director and daily tester. Claude Code wrote nearly all of the code. My job was to set direction, review every build on my desk and my phone, and make the calls.

The project started from a technical thesis: an agent that reads measurements instead of pixels will model better. Every face and edge has a stable ID, the renderer reports its shading in parts, and the document behaves like a program. The first version looked like a well-behaved Mac app: an outliner, a Metal viewport, an inspector, a modifier stack, a progressive path tracer, and an Assistant panel built into the sidebar.

125
Agent tools
108 editing, 17 reading, from one tool table
14
Named releases
4 previews, then 10 betas
3.6 MB
Modelling kernel in the browser
Same engine as the Mac app
Early dark-mode Mac app with a floating bottom toolbar and an empty render pane reading: Render to see the scene as light actually finds it

An early redesign: dark UI, a floating toolbar, and the render docked as a split pane instead of a separate window. The empty state explains what it does.

The test object was a LEGO 2×2 brick. It's mostly constraint (stud pitch, wall thickness, tube diameter), so whether it was right was a fact, not a taste call. It surfaced problems fast, including an audit that found seventeen of the twenty-seven material controls did nothing.

Clean glossy red 2x2 LEGO-style brick render

The brick, built from a script with creases that hold. Earlier attempts came out looking like marshmallow.

Losing to Blender

Test your thesis early enough that you can afford to lose it

Instead of arguing about whether the idea was good, we ran trials. Fresh agents with no memory of the project got a numeric spec and a tool, and a harness logged every command.

The first finding had nothing to do with the thesis: 26 to 35% of agent commands failed because the agent couldn't find out what arguments a tool took. One new tool, describe_tool, brought that under 1%. It was the agent-facing version of a discoverability fix, and the cheapest big win of the project.

Then the real test. Agents built the same part to a ten-point spec in Blender and in Lingisa:

How the agent workedPassedTimeTokens
Blender, through its Python API5/5169 s52k
Lingisa, through its own command vocabulary5/5543 s104k
Lingisa, through a new Python binding5/5162 s56k

Lingisa could match Blender, but it couldn't beat it at modelling alone. "Hand back measurements rather than pictures" wasn't a differentiator against a tool with a Python API. The lesson we wrote down: an agent should drive the tool in a language it already knows. Never invent a language.

The rewritten mission

A person and an AI work on the same model, in the same document, and hand it back and forth. The person points at the part they mean and says what they want; the agent does it to that object, in that place, says what it did, and leaves anything it guessed at marked for a person's eye.

What survived wasn't modelling power, it was collaboration. Almost all of the design work that follows came from that reframe.

Designing the Hand-off

Everything that shipped next followed from the mission: comments pinned to faces and edges that the agent reads first, a watcher that picks them up, variants kept in place, one undo per agent request, a visible agent with a Stop button, and an acknowledgement within two seconds.

Blue 2x3 brick with a side hook

Built entirely through comments. I asked for a 1×1 brick, then inner ribs, then a hook. Each reply said what changed, in millimetres.

The agent's word isn't the last word

When the agent finished, its comments marked themselves resolved. That felt wrong: the person asked, so the person decides. Done work now waits in "awaiting review" until someone chooses Accept or Not quite…

A comment is a note until it @-mentions someone

Only comments addressed to the agent get acted on. People can talk to each other in the margin without the agent jumping in.

"On it."

The agent worked so fast that I didn't notice the progress banner. The banner now starts with an acknowledgement. Two seconds of "On it." changed how it felt more than any speed-up did. What a person feels is the hand-off, not the compute.

Variants: the hat on the head

The obvious way to show AI options is a row of copies. But if you were working on a hat on someone's head, you wouldn't make three next to each other. You'd switch between versions on the head. So a variant is a look kept in place, switched with arrows on the object itself, and a colour exploration and a shape exploration became the same feature.

Comments panel full of resolved threads next to a wireframe character and its path-traced preview

Split view: the wireframe on the left, the traced preview on the right, and the thread with the agent alongside. Each comment is for someone, and done work waits for review.

Variants shown as large cards, each traced in a studio chosen by its material

Variants started as a list of tiny thumbnails. Now they're large cards, and each is traced in a studio picked by its material: light grey for dark tones, dark for pale colours and glass.

Making Light Honest

Pictures sell the work, so the renderer got real attention. We added a denoiser, then measured it against ground truth rather than eyeballing it. At first it destroyed 73% of the detail seen through glass. After the fix it kept 96%, and 16 denoised samples beat 128 raw ones. HDRI lighting, depth of field, a shadow catcher and transparent backgrounds followed, so a render can drop straight into a layout or sit inside a photo.

Denoised render of a brick scene
Glass block sitting in a grassy field panorama

Left: denoised render, checked against ground truth. Right: shadow catcher plus panorama lighting, placing an object into a photo.

Materials got the same scrutiny. The original glass controls asked for per-metre absorption coefficients, which means "pick red to get blue." I replaced them with Seen through, Tint and Cloud: describe glass the way you see it. The agent's tools still take the physics.

Three-view reference drawing of a light aircraft
Rendered white light aircraft with a red tail

A three-view reference handed to the agent, and the aircraft it built from it.

From Mac to Browser

The original web plan was "a prompter and a viewer, not an editor." Editing would stay on the Mac. Two things changed that. First, reach became the point: someone curious should be able to open a link and explore without installing anything. Second, the expensive part turned out to be cheap. A WebGPU path-tracer spike ran faster in Chrome than the native Metal tracer on the same Mac, and the modelling kernel compiled to WebAssembly with all but one test passing.

We reversed the decision, and kept the old reasoning at the top of the plan rather than deleting it. The one constraint that didn't move: one kernel. The web app is a React shell over the same engine, not a second modeller. Figma's architecture was the reference: an engine in WebAssembly, a React interface, its own GPU renderer, and a server that orders edits.

The browser version quickly got things the Mac didn't: sign-in, a documents list, per-person selection and presence, a materials palette, and Mark a space. You mark a region where something might go, and the agent fills it with ghost proposals that stay out of shadows and renders until you keep one. Keeping a proposal is the approval.

Documents page with path-traced thumbnails for each document

Documents with traced thumbnails. Before one is traced, a card shows an isometric block in the document's own colour.

Rendered sports car wheel with a yellow brake caliper

A wheel built by an agent working over the web. Agents filed bugs as they worked, which meant the product thesis was improving its own development.

Simple Mode

The biggest simplification of the project

Eventually the panels outgrew the mission. A material showed 25 controls. That serves someone who models, but Lingisa is for someone curious who opens a link and lets their agent do the work. So panels now start Simple: a handful of controls, plus one-tap asks like Frosted glass, Round the edges or Softer, and a free-text line. Each ask becomes a comment for the agent, and you watch its steps appear under it.

Before

25 material controls, slider-heavy, built for a 3D artist. A new visitor had no idea where to start.

After

A few controls plus one-tap asks. A Full controls switch brings every setting back, and every setting is still a tool, so an agent can reach all of them anyway.

Lingisa in Simple mode with a brick selected, one-tap asks like Round the edges and More detail, an Ask Agent field, and a Getting started card

Simple mode: transform, four one-tap asks, a field to ask the agent, variants, and a Getting started card written for a first visit. The toolbar leads with Mark a space.

Bring Your Own Agent

The first version had its own built-in assistant that asked you to paste an API key into a dotfile. It was the first big thing I cut. It duplicated what external agents already did with a real context window and the full tool table. The app's job became being the best possible place for any agent to work.

Lingisa doesn't ship its own AI. You connect yours: Claude Code, Codex, Gemini CLI, Cursor, VS Code, ChatGPT or the Claude app. A small npm connector delivers a comment to your agent in under 200 ms. Setup is one dialog with one command to copy and one sentence to say to your agent. Copy is agent-neutral throughout: "Ask Agent," "Agent connected."

Connect your agent dialog with an agent dropdown and connection scope
Comment thread where the agent explains its lighting changes

Connecting an agent, and a reply where it explains why area lights failed on glass, what it tried instead, and how to adjust it yourself.

A mascot exploration shows the full loop. I asked for "more creative solutions, directions we haven't explored." The agent built concept rows in the scene beside the original, labelled each one, and kept body options as variants I could flip through on the object. Its reply names each concept by position and links each part, so clicking a name selects it. Nothing generated was thrown away.

Agent reply describing three rounder mascot concepts next to a wireframe robot character
A lineup of robot and astronaut mascot concepts rendered together

The whole exploration in one render.

Small Decisions

Where I simplified, rethought, and finessed

No sliders

Every number became a scrubbable label plus a typed field, with modifier keys for coarse and fine. A slider can't express 0.1471, but it looked like the only control.

The camera is not the viewport

My note was "I feel like the camera doesn't do anything." It turned out the camera was the viewport, so any orbit moved your shot. I split a free view from cameras that hold a framing. New documents open in the free view, so moving is safe from the first second.

Selection belongs to the viewer

Each person has their own picks and view. Your selection is never someone else's undo step.

Saved has three states, not two

Green with a time, amber, and yellow "Not in a file." Saying "Saved" about a recovery file would tell someone their work is somewhere it isn't.

A tool UI at Figma's scale

Early builds used 13px text and 36px rows, which felt large for a tool. I moved to Figma's scale (11px labels, 28px rows), rebuilt the stylesheet on design tokens, and kept 36px rows and 44px touch targets on phones.

Accessibility from the start

Full keyboard use, an interface size setting from 90 to 200%, single-key shortcuts that can be switched off, reduced motion, 4.5:1 contrast, and automated accessibility checks in CI.

What I Learned

This was one designer and one AI pair-programmer. The working agreements were written down: an ordered task list where every item had a stated reason, a history file with the write-up for every decision (including retractions), and a proposals list where nothing started until I marked it approved. A build number sat on screen at all times, so I always knew whether I was reviewing the new build or a cached one.

The most useful habit was keeping the wrong turns in writing. The original thesis lost. A feature built as an A/B experiment showed no measurable effect. Given a deliberately wrong spec, agents bent the model to fit it rather than questioning it. Each of those is still in the record, because the mistake is the useful part.

What I'd tell another designer

Test your thesis early enough that you can afford to lose it. Losing to Blender gave us the actual product.

Next on the roadmap: reviewing together, with viewers and suggestions that work like a pull request on a model; product shots, with the renderer as the end product; and more modelling power once the collaboration layer is solid.