Skip to content

A Prompt Is Not a Request, It Is a Work Specification

Adityo Guni Waluyo

The same agent task produced different results every run. The prompt left decisions to the model; the fix is an explicit work specification.

TL;DR

Same prompt, different outputs usually means the prompt underdefined the work, not that the model is unstable. Treat prompts as work specifications: state objective, context, constraints, output shape, and a verification step. Clear specs make runs converge, checking objective, and cut the expensive corrective iterations, regardless of which model you use.

The same agent task, say writing a README for a new module or generating a data validator, produced different results on every run. The prompt stayed exactly the same. The first guess blamed the model: temperature too high, or the model simply unstable.

The execution logs said otherwise. The prompt underdefined the work. The model silently picked default values for every aspect left unstated, and each run picked differently. The problem was not the model; it was an instruction that left too many decisions to the model.

The mental model then shifted: a prompt stops being a request and becomes a work specification. A recent academic tutorial frames prompt engineering as the discipline of turning informal human intent into structured AI work specifications, and states plainly that the strongest prompt is rarely the longest; it is the one that makes desired behavior, required sources, and success criteria unmistakable.

The minimal binding specification

A structure proven sufficient for daily tasks: objective, context, constraints, output shape, and a verification step. An empirical study at ICPC 2026 elicited ten guidelines for improving code-generation prompts, and they all revolve around the same moves: specify input and output, give pre- and post-conditions, provide examples, resolve ambiguities. Post-conditions deserve a close look because they help the model verify the correctness of its own output logic before handing it to a user.

At the organizational level, the same pattern appears as policy. Standard PRD-STD-001 requires production prompts to carry an objective, context, and explicit constraints, while forbidding credentials or personal data inside instructions. The rationale is measurable: ad-hoc prompting leads to inconsistent outputs, higher defect rates, and wasted iteration cycles.

In daily workflow this can be as small as changing how you ask. Instead of "write a validator for the user form", the specification reads: write a validation function for a registration form; context: a city portal web app; constraints: use TypeScript, reject emails without a valid domain, return a structured error object; verification: include three test cases, one with empty input.

What changed after the standard landed

Run results started converging. Outputs matched in shape across runs, and checking became objective: success criteria were written into the prompt instead of living in a developer's head. Spending a few extra minutes framing context and constraints up front removed the repeated corrective iterations, and that is the expensive part.

A standard like this also survives model churn. Models will keep changing, but a good work specification keeps expected behavior, boundaries, and how to check them explicit. The three sources above, from an academic tutorial, an empirical study, and an organizational standard, converge on the same point: specification clarity determines the outcome more than the intelligence of the model used.

Sources

A Tutorial on Prompt Engineering: From Messy Thoughts to AI Workflows, accessed October 10, 2026.
Guidelines to Prompt Large Language Models for Code Generation, ICPC 2026, accessed October 10, 2026.
PRD-STD-001: Prompt Engineering Standards, accessed October 10, 2026.

Related articles