An AI demo shows that a model can produce a useful result for a prepared example. Production asks a harder question: can the complete workflow run repeatedly, within its agreed risk and operating constraints, and at an acceptable cost when inputs are messy, APIs fail, and the original builder is unavailable?
This simplified example is informed by the crawl-to-publish workflow described in the WebLingo case study. It is not a customer case study and makes no claim about customer outcomes.
TL;DR: Production readiness is a decision about the whole AI workflow, including permissions, failure handling, cost and operational ownership. Use evidence to decide whether to continue, narrow the scope, redesign or stop. The illustrative checklist helps organize that decision; it is not a certification.
This guide uses a simplified website-localization workflow to explain eight production-readiness checks: outcome, system boundary, evaluation, permissions, failure recovery, cost, latency and capacity, observability, and rollout.
Here, production-ready means that the scoped workflow meets its agreed outcome, risk tolerance, and operating constraints. It does not mean that a model is universally safe or that every future use case has been solved.
What an AI demo shows and what it leaves unanswered
Assume a team wants to use an AI system to turn a website source page into a draft localized page. A reviewer should be able to inspect the draft, correct it, approve it, and publish it at the intended URL. The demo shows a translated paragraph; the website-localization workflow needs a source record, locale and glossary context, reviewer permissions, a publish decision, and a recovery path.
| Readiness area | Question to answer | Example evidence to collect |
|---|---|---|
| 1. Outcome | What useful result must the workflow deliver, and for whom? | Defined review and publication outcome; approval or rejection recorded |
| 2. Data and system boundary | What data enters each component, who can access it, and what can leave the system? | Source page, glossary, locale, provider access, storage, and retention map |
| 3. Evaluation | How will quality be tested repeatedly rather than judged from memory? | Representative pages, error rubric, glossary checks, reviewer notes, and regression cases |
| 4. Security and permissions | Can the workflow propose or perform an action beyond the user's authority? | Role-based review and publish permissions, least-privilege credentials, and audit record |
| 5. Reliability | What happens on timeout, invalid output, duplicate processing, replay, or retried delivery? | Schema validation, retry limits, idempotency, fallback, and recovery procedure |
| 6. Cost and latency | What cost, response-time, and throughput budgets are acceptable? | Model cost per approved page, review time, queue depth, end-to-end latency, provider quota, sustainable throughput, and retry volume |
| 7. Observability and ownership | Can an operator understand what happened without asking the original developer? | Logs, model and prompt versions, alerts, escalation owner, and support procedure |
| 8. Rollout and rollback | How will exposure increase, and which conditions stop or reverse the release? | Internal use, limited pages or locales, monitored expansion, and versioned rollback |
The table is a diagnostic prompt, not a certification checklist. A scoped workflow may need more or less evidence.
Eight production-readiness checks
1. Define the outcome and acceptance criteria
"Use AI to localize the site" is not an acceptance criterion. Name the user, the action, the allowed data, and the decision that follows. For this example, a reviewer approves or rejects a generated page before publication, and the system records that decision.
2. Map the data and system boundary
List the source system, model call, storage, reviewer interface, delivery host, and owner for each handoff. Record which data may be sent to the approved model provider, where it is retained, and who can access it. The NIST Generative AI Profile is a useful lifecycle reference for this context and risk work.
3. Build a representative evaluation set
An evaluation is a repeatable way to compare an output with an expected result or rubric. Start with representative and difficult inputs, define what counts as wrong, and keep production failures as regression cases. OpenAI's evaluation best practices describe one provider-specific approach; the method should remain understandable if the model or vendor changes.
4. Restrict permissions and side effects
Security and output validation are separate gates. Decide which data may leave the system, who can invoke the workflow, which actions the model may propose versus perform, and whether tool calls inherit the user's permissions. Validate structured output against a schema, sanitize or escape generated HTML for its rendering context, and authorize every side effect independently of the model output. Tools should run with scoped, least-privilege credentials rather than unrestricted user or application access. The OWASP LLM Top 10 is a risk taxonomy, not a full security audit.
5. Design failure and recovery paths
Plan for timeouts, rate limits, malformed output, duplicate processing, replay, retried delivery, and partial publication. Validate the schema before a side effect, cap retries, make retries idempotent, and give a person a clear recovery path. In the localization example, a failed page should stay unpublished while the reviewer can retry or restore the last approved version.
6. Set cost, latency, and capacity budgets
Cost, latency, and capacity are production constraints, not proposal footnotes. Track model cost per approved or published page, human review time, queue depth, end-to-end processing time, provider quotas, sustainable throughput, publish failure rate, retry volume, and the cost of duplicate work. A workflow that is affordable and fast for ten pages may still fail at ten thousand pages or when human review becomes the bottleneck. The acceptable numbers depend on the workflow's value and risk; agree the budget before expanding exposure.
7. Add observability and operational ownership
Log the input and output identifiers, model and prompt versions, validation result, reviewer decision, publish action, and recovery event. Define the operator, escalation owner, alert threshold, and support procedure so someone can understand what happened without asking the original developer. This is where who should own the deployment becomes an operational decision, not a job-title debate.
8. Roll out gradually with stop conditions
Write down which quality, latency, cost, or failure signals pause expansion. A rollout is more controlled and reversible when its exposure and stop conditions are explicit; a named owner must accept the residual risk before expansion.
- Offline evaluation. Run representative and difficult cases without affecting users.
- Shadow mode. Process real inputs without publishing or triggering external actions.
- Internal release. Let only the team review and approve results.
- Limited production release. Enable one locale, workflow, customer group, or content type.
- Measured expansion. Increase exposure only while agreed quality, latency, cost, and failure metrics stay within budget.
- Rollback or pause. Stop on a defined failure and restore the previous version.
Worked example: a website-localization workflow
The useful path is not "prompt in, page out." It is a chain of decisions and evidence:
- Source
Capture content, source version, and route.
- Context
Attach locale, glossary, and source metadata.
- Authorize
Enforce reviewer and publisher permissions in application code.
- Generate
Produce a versioned draft.
- Validate
Validate structure, required terminology, and publication rules.
- Review
Approve or reject with notes.
- Publish
Publish the approved version, monitor it, and retain a rollback path.
Useful measures include cost per approved or published page, percentage approved without substantive edits, review time per page, queue depth, end-to-end latency, provider quota, sustainable throughput, retry volume, and publish failure rate. Two practical failure cases are a required glossary term translated incorrectly, causing reviewer rejection, and a timeout after validation but before publication; both should be visible, recoverable, and included in later regression tests.
| Decision | Illustrative rule | Evidence |
|---|---|---|
| Scope | Begin with one locale and selected page types. | Page and route inventory |
| Publication gate | Only a reviewer-approved version may go live. | Approval record and version identifier |
| Quality gate | Required fields, glossary terminology, and route checks must pass. | Validation and reviewer results |
| Failure rule | Invalid output or a timeout cannot replace the currently approved page. | Job log and recovery test |
| Expansion | Add pages or locales only while quality, latency, and cost stay within budget. | Rollout review |
| Stop condition | Pause on an unapproved publication, unrecoverable route conflict, or sustained budget breach. | Alert and incident record |
When to continue, narrow, redesign, or stop
These four outcomes keep a readiness review from collapsing into "ship" or "do not ship."
| Decision | Use it when | Next action |
|---|---|---|
| Continue | Critical checks remain inside the agreed quality, risk, and operating budgets. | Expand by one rollout stage. |
| Narrow | One locale, page type, user group, or action is ready, but the broader scope is not. | Reduce the deployment boundary. |
| Redesign | Failures, permissions, evidence, or ownership cannot be isolated clearly. | Change the workflow or architecture and retest. |
| Stop | Expected value no longer justifies the risk, cost, or operating burden. | Close the pilot and document the findings. |
Frequently asked questions
What makes an AI prototype production-ready?
The scoped workflow meets its agreed outcome, risk tolerance, and operating constraints, with evidence that someone else can run and support it.
Can a system be production-ready if every output needs human review?
Yes. Human review is compatible with production when the rubric, permissions, logging, escalation path, and reviewer capacity are explicit. The problem is ad hoc judgment by the original builder.
How many examples should an AI evaluation set contain?
There is no universal number. Start with representative, difficult, and known-failure cases; add production failures as regression cases; and expand until the set covers the decisions and risks in scope.
What should be monitored after deployment?
Monitor quality signals, latency, cost, retries, failure rates, reviewer decisions, permission events, and rollback triggers. The right thresholds depend on the workflow's value and risk.
When should an AI pilot be stopped rather than expanded?
Pause when agreed quality, cost, latency, risk, or recovery thresholds are breached and the team cannot restore them inside the current boundary.
Sources and further reading
- NIST Generative AI Profile, a companion risk-management reference for generative AI.
- OpenAI evaluation best practices, a provider-specific example of repeatable evaluation work.
- OWASP LLM Top 10, a security-risk taxonomy for generative AI applications.