Check LLM output in code, then retry exactly once

  • ai
  • llm
  • testing

I run a small side project that turns news stories into short articles for people learning French, at several reading levels. A language model does the writing. Every story gets a headline per level, and the prompt has always said, in plain words, “80 characters maximum.”

When I actually measured a few days of output, about one headline in five was over 80. Some were close (82, 83), some were not (90). The model was not ignoring the instruction so much as treating it as a suggestion, which is exactly what an instruction in a prompt is.

The fix was not a better prompt. It was a pattern I now use for almost every constraint I care about in generated text: check it in code, feed the problems back for one rewrite, and decide ahead of time what happens if the problem survives.

A prompt is a wish, a check is a rule

If a limit only exists as prompt text, nothing knows whether it held. You find out when a reader, a layout, or a search snippet finds out for you. So the first step is boring on purpose: write a pure function that measures the output and returns a list of problems.

const TITLE_MAX_CHARS = 80; // the same number the prompt states

function titleLengthProblems(titles) {
	const problems = [];
	for (const [level, title] of Object.entries(titles)) {
		const n = String(title ?? '').trim().length;
		if (n > TITLE_MAX_CHARS) {
			problems.push({
				where: `title ${level}`,
				type: 'title too long',
				message: `"${title}" is ${n} characters; the limit is ${TITLE_MAX_CHARS}. Shorten it and keep exactly the same facts.`,
			});
		}
	}
	return problems;
}

A few details here matter more than they look:

  • It only measures. No LLM call, no side effects, so it is trivial to unit test and costs nothing to run on every generation.
  • The problem shape matches every other check. Length, grammar, and “does this contradict the source” all return { where, type, message }. That lets them share one list and one rewrite loop instead of each growing its own.
  • The message is written for the model, not for me. It quotes the whole title, gives the actual length and the limit, and says what to preserve. Vague feedback (“too long”) gets vague fixes.
  • The constant points at the prompt. The number lives in two places (the prompt and the check), so a comment ties them together. If one changes, the other should too.

One rewrite, not a loop

With problems in hand, the generator gets exactly one chance to fix them:

async function generateWithCheck(generate, check) {
	let output = await generate();
	let problems = await check(output);

	if (problems.length) {
		const feedback = problems.map((p) => `${p.where}: ${p.message}`).join('\n');
		output = await generate(feedback);
		problems = await check(output);
	}

	return { output, problems };
}

Why once? Because a second and third retry rarely buy what they cost. If the model did not fix a clearly described problem with clear feedback, asking again mostly burns tokens and latency, and an unbounded loop is how a scheduled job ends up running for an hour on one stubborn input. One rewrite catches the common case (the model just overshot) and keeps the worst case bounded and predictable: at most two generations, always.

It also makes the behavior easy to test. I inject generate and check as fakes and assert the call count: one generation when everything passes, two when something fails, and never three.

Decide what a surviving failure means

This is the part I see skipped most often. After the one rewrite, some problems will still be there. What you do with them should be a deliberate choice per problem type, not whatever the code happens to do.

In my pipeline, leftover problems normally become gate failures: the story is saved as a draft for a human instead of being published. That is right for things like a headline that contradicts the source. A wrong fact is worse than a missing story.

A headline that is 84 characters is not in that category. Holding an otherwise good story back because its title is four characters long would be a strange trade. So length gets a different outcome:

const { output, problems } = await generateWithCheck(generate, check);

const tooLong = problems.filter((p) => p.type === 'title too long');
if (tooLong.length) {
	log(`WARNING: ${tooLong.length} title(s) still over the limit after the rewrite`);
}

const blocking = problems.filter((p) => p.type !== 'title too long');
// blocking problems hold the story as a draft; warnings do not

Same check, same rewrite, different consequence. The warning still lands in the log, so I can count how often it happens.

Don’t “fix” it with string slicing

The tempting shortcut is title.slice(0, 77) + '...'. I deliberately do not do that. A truncated headline can drop the one word that makes it true (“Minister denies resigning” cut down to “Minister…resigning”). Code is good at measuring text and bad at editing its meaning. If the text needs to change, the model should change it, with the facts it was told to keep. If it cannot, a slightly long but correct title beats a short wrong one.

Treat it as a hypothesis, not a fix

The last habit is the one that keeps this honest. Before shipping the check, I wrote down the baseline (5 of 24 recent titles over the limit) and what would count as working two weeks later: zero over the limit, and no increase in stories held as drafts. If long titles keep showing up at the same rate, that is a real answer too. It means this model does not respond to length feedback on a single rewrite, and the fix has to be a different mechanism, not more retries.

There is also one honest risk to watch. The rewrite regenerates all the titles, so a story whose only problem was length can come back with a new, blocking problem it did not have before. That is the mechanism working as designed, but it is the one path where this change can cost a story, so it is the first thing I look for in the logs.

The pattern, in short

  1. Every constraint you care about gets a check in code. If it only lives in the prompt, assume it is broken some percentage of the time.
  2. Checks are pure and return one shared problem shape, with feedback written for the model.
  3. Retry exactly once with that feedback. Bound the worst case.
  4. Decide per problem type what survives: block, warn, or accept. Write it down.
  5. Never repair meaning with string operations.
  6. Measure before and after, so you know whether the check actually changed anything.

None of this is specific to headlines or to French. It works for JSON that has to match a schema, summaries that must stay under a word count, or code that has to compile. The model does the creative part; the code decides what counts as done.