SH Scott Hantler

Writing · · v1.0

The Wrong Question

AI can answer the question, build the thing, and verify the result. None of that means the question was right.

I've had this URL for a long time. I bought it back when buying URLs was a thing you did, just to save the name, and did nothing with it.

Then there was an event for an initiative I work on, less than a week out, and I needed something to hand somebody. I didn't have a business card. You don't hand out a resume and I wasn't going to say look me up on LinkedIn. I needed a calling card. Something that said me, this is what I did, this is why I'm here.

I had nothing that said me. A resume says where I worked. LinkedIn lists the organizations. An elevator pitch is the wrong thing entirely at a museum opening.

So I asked an AI, Claude, to build me a website.

Then I asked a different AI, ChatGPT, what was wrong with it.

It came back with the right diagnosis and it came back fast. The site took too long to say who I was. A stranger landing on it needed one plain sentence, my name and what I do and why I'm there, before anything else. The model wrote the sentence. I agreed with it. I used it.

I just didn't put it at the top, where it belonged. I buried it in the middle of a paragraph where nobody was going to find it.

The advice was right and I took the advice and used the sentence. And the outcome was still wrong, because I'd asked what sentence I was missing, and I never asked where the sentence should go.

I know that happened at 5:02 on the afternoon of Wednesday, May 20, and I know it because the same AIs that built the site went back through the record and dug the timestamp out for me while I was writing this. Hold onto that, I'll come back to it.

Four hours and forty-six minutes later the first commit landed, the first save point in the build's permanent record. Ten files, 935 lines added, and most of the ideas that are still holding this page up as you read it. Nobody would have blamed me for a placeholder that weekend, the coming-soon page you put up to hold an address. What was standing by 9:48 that night was the real thing.

The models generated the site's custom code. I directed the work, selected what survived, and accepted the result; I never touched the HTML or CSS myself.

There's a strange problem in how we talk about this, and I haven't found a satisfying way to describe it. We keep trying to locate the authorship, the artistry, in the keystrokes: the code, the prompts, the ones and zeros. If a person typed it, the person wrote it. If a model generated it, the model made it. But where's the fuzzy line between pure generation and what a person can honestly claim as ingenuity, as ownership?

My first job was on TV and film sets. On Law & Order, I was a production assistant: the gofer, the kid who ran the coffee, blocked doors, and waited through lockup while each department closed out its work.

I didn't know every craft in detail. What I learned was how the handoffs worked. A camera operator, a cinematographer, and a director could all be shaping the same scene without doing the same job.

The work overlapped. Responsibility did not disappear.

Back in May the code came fast and the site went up. I had a calling card ready before the doors opened. What didn't come automatically was the answer to a much more basic question, which was what the thing was supposed to do, and why I was doing it at all.

I thought I knew. I needed a site by Saturday. That was true and it wasn't enough.

Here's the humbling part, the one I don't love making public.

Because it was my own site and every change was easy to reverse, I let myself move without the reviews and close reads I would have insisted on if the work had been for someone else.

I mistook cheap iteration for low stakes. The distinction mattered.

That turns out to be the whole story. I just didn't know it yet.

The event went well, by the way.

The site, the thing the whole rush had been for, never came up. I don't think I mentioned it to a single person I met.

Four days later, on another whim, I built five versions of the transition that plays when the page changes and discarded four. The record moves through a column scan, a dimensional flip, token rain, and what the machine called a WebGL tear. I couldn't have told you then what WebGL meant. I did know this site was not supposed to become a gimmick. By evening the effect had been rebuilt around the actual rendered page.

That's a lot of attention for a transition. I know. It's also where the division of labor got visible to me for the first time.

A model can make five transitions. It can make fifty. It can compare them against a brief, check the code, explain the tradeoffs, and hand back more before I have decided what any of them should feel like.

What none of them told me was that the fourth one felt like fireworks, and that fireworks were wrong for what the page was supposed to become. That was not a code defect. It was a judgment about the page.

The effect wasn't broken, and the AI hadn't done anything except exactly what I asked. It was doing the wrong job because I'd given it the wrong job.

AI made the production fast. I saved each version before the next one touched it and wrote down why I discarded it, so going backward cost almost nothing. Trying something novel stopped being a bet I had to win.

A wrong answer, recorded, was still an answer.

Somewhere in there, the commit history stopped being only a record of code and became a record of decisions. Reading back through the messages now is like reading early drafts of this essay. One records why an effect stayed: without it, there would be no visible sign that anything was happening during the page change. Another, from the night I took the impact numbers off the site, records the line I drew. The claims came off. The stop counts, section numbers, and tax category stayed on the page because a reader could still use them.

That's where my hand is clearest: in what I asked for, what I kept, what I cut, and why. The machine produced the lines. The decisions were mine.

There's a conversation across AI about keeping a human in the loop. In practice the phrase covers many arrangements: feedback, review, intervention, final decision. When I started using ChatGPT, it often felt literal: one prompt in, one answer out, the person carrying the answer from one window to the next. Now a system can keep working after the prompt. I understand the handbrake image, the person bolted onto an assembly line so it cannot run past a failsafe. But that wasn't the most important role I was playing on this website.

The machines could produce the substance of that page faster than I could use it. What I had not asked them to do, and what they could not be accountable for, was decide how the page should feel to someone arriving cold rather than someone I'd invited, how credit should be assigned, or whether a genuinely impressive effect had quietly turned the whole page into a gimmick. They weren't failing to have taste. They were doing the work I'd given them. What stayed mine was deciding whether I'd given them the right work at all.

I didn't understand yet how literal that was about to get. Not until I tried to write this.

In July I started the site over. When it was up again I put it in front of five separately prompted AI readers, cold, with no project history and nothing but the page. I showed it to people too, humans whose judgment and taste I trust.

They converged on the same question about what I had actually done. Strip the politeness and it's this:

Cool project. What'd you do for it?

Seven weeks earlier the very first AI I ever asked to review this had told me the exact same thing. The site took too long to say who I was. It had written the line that would fix it. I put the line in the wrong place and moved on.

The first answer was right on day one, and right hadn't been enough.

What changed the second time wasn't the information. It was the arrangement of it. The first time it came as one bullet in a list, and I treated it like just another bullet in a list. The second time, the machines and the people all came back with the same thing, without having compared notes.

So I moved the sentence to the top. My name is there now, above the project, and above this essay.

Six days later I went back to the machines and asked them to reconstruct how the site had been built.

By then the record was much bigger than the site. Transcripts, reviews, the code, the commit history, deployment records, dead ends, and the kind of thing a person actually types at a machine when the week is long and the thing has to work. One message from the last day of that first week reads, in full: "Done. Try again. Launch."

As with most of my work, I wanted to know how this project had turned out, so I ran an evaluation and a post-mortem. I wanted a bounded account of the founding week.

On July 15, working with an AI, I approved a spec that told the systems which dates to look at. The dates I gave were wrong. The search window opened six weeks after the period it was supposed to recover.

Everything downstream looked excellent. Claude ran the analysis. ChatGPT checked it. Local scripts re-derived the counts. The reports agreed on the files, the comparisons returned no mismatches, and a regression check confirmed that the window was the only thing that had changed.

Every check passed and the answer was worthless.

This was not a hallucination or a miscount. The probabilistic systems produced and checked the analysis; deterministic scripts re-ran the counts. Together they established that the answer followed faithfully from the window, rules, and inputs I'd supplied.

The machinery worked exactly as specified, which is why it took me so long to see that the inquiry had failed.

What exposed the failure was something I remembered.

The analysis said there were no site reviews in the window. I knew one existed. I'd read it. So zero wasn't surprising, zero was impossible, which meant either my memory was wrong or the frame was.

I checked the frame.

The window was wrong because the specification was wrong, and the specification was wrong because I approved it without the scrutiny it needed. Three separate checks then confirmed that the execution was faithful while sharing the same premise. I had supplied the wrong frame, and every check after that was rigorous about the wrong thing.

Verification can tell you whether an answer follows from its inputs. Validation asks whether those inputs describe the question you meant to ask. Every check inherited the same search window. The checks were separate. Not one was independent of the premise.

You could ask a model to challenge the premise. Of course you could. Require a date sanity check, commission a search designed to contradict, tell one system to assume the spec is wrong. I will next time, and I'll build it into my automated workflow. But that only moves the question back a level. Challenge what? Which assumptions get suspected? What fact from outside the frame would let it know that zero is absurd rather than merely zero?

You can build machinery that asks more questions. Somebody still has to know when zero is impossible. In this case it was a memory I'd never put into the system. Everything I gave that system, the system could check. What caught the mistake was the one thing outside the frame.

And here's the part I keep thinking about. The conversation the dig couldn't find was the one from May 20. The review that told me, before there was a single line of code, that the site took too long to say who I was. It had been sitting in the archive the whole time. It scored zero because I wasn't searching the right place.

That 5:02 I gave you at the top came out of that same dig, once it was pointed at the right two weeks.

Some of the loudest conversations about AI risk are about the machine getting out of the box. That risk is real. OpenAI disclosed that models in an internal cybersecurity evaluation, operating with reduced safeguards, had found a path to the open internet and compromised Hugging Face's production infrastructure in pursuit of benchmark solutions.

Days later, Anthropic disclosed a different failure: Claude models were told they had no internet access, but a misconfiguration left an open path, and they gained unauthorized access to real systems at three organizations while initially treating those systems as part of the exercise.

Mine was different. I had approved the wrong search window.

The machinery did exactly what I asked. I had asked the wrong question. Three separate checks inherited the same window; every check came back green, and the answer was worthless. I caught it because one fact didn't add up.

There is a quieter risk too. These tools keep getting better at doing more without stopping. More of what once required a handoff and another approval now happens inside the loop. Most of the time, the work is good. That's what makes the wrong answer easier to miss.

The pause used to be part of the work. Now it's a setting.

Always allow. Don't ask again.

That's how a near miss becomes a failure: you accept the answer and move past the one thing that doesn't add up.

I thought I had guarded against it. That's why I ran checkpoints, cold reads, and separate checks.

None of them questioned the window I had approved. They checked the work, not the question.

The wrong question I asked was about dates: two weeks in July instead of two weeks in May. Breaking the rules is not the only way this goes wrong.

Sometimes the machine follows a bad premise exactly. Every check passes. Every verifier agrees. You can get something wrong much faster than it used to take you to be wrong. The answer can look finished and the question can still be wrong.

I've been making the same claim about other people's work for years: the substance of a project is rarely what holds it back. The harder part is its shape. I wrote it about projects that were real and still couldn't get into the right room.

Then I spent two months proving it on myself.

My first job had a name for the rest of it: the part between the judgment and the code. Since I started working with these machines, I've been the gofer again, walking every file, every spec, every answer from machine to machine, deciding what each one needs to see, and remembering which conversation holds the current truth.

I keep track of the handoffs because otherwise speed only multiplies confusion. The machines did the typing and multiplied the carrying. Nobody warns you about that part, and no demo ever shows it.

What I'm building now is the production office: the version of this work where the carrying stops being mine and the deciding stays mine.

I didn't delegate the judgment. I delegated enough of the work to find out where the judgment actually was.

Process note

My wife has edited my writing since high school. I accept some edits, reject others, and rewrite the ones that improve the work but stop sounding like me. She, along with Claude, Gemini, and ChatGPT, assisted here with reconstruction, checking, red-teaming, and editing. The final text, its factual claims, and its judgments are my responsibility.