AI deployment

AI Demo to Production: Readiness Checklist

Move an AI demo toward production with an illustrative, practical checklist for evaluation, security, failure handling, cost, ownership, and staged rollout.

François Guéguen 9 min read

An AI demo shows that a model can produce a useful result for a prepared example. Production asks a harder question: can the complete workflow run repeatedly, within its agreed risk and operating constraints, and at an acceptable cost when inputs are messy, APIs fail, and the original builder is unavailable?

Illustrative example, not client work

This simplified example is informed by the crawl-to-publish workflow described in the WebLingo case study. It is not a customer case study and makes no claim about customer outcomes.

TL;DR: Production readiness is a decision about the whole AI workflow, including permissions, failure handling, cost and operational ownership. Use evidence to decide whether to continue, narrow the scope, redesign or stop. The illustrative checklist helps organize that decision; it is not a certification.


This guide uses a simplified website-localization workflow to explain eight production-readiness checks: outcome, system boundary, evaluation, permissions, failure recovery, cost, latency and capacity, observability, and rollout.

Here, production-ready means that the scoped workflow meets its agreed outcome, risk tolerance, and operating constraints. It does not mean that a model is universally safe or that every future use case has been solved.

What an AI demo shows and what it leaves unanswered

Assume a team wants to use an AI system to turn a website source page into a draft localized page. A reviewer should be able to inspect the draft, correct it, approve it, and publish it at the intended URL. The demo shows a translated paragraph; the website-localization workflow needs a source record, locale and glossary context, reviewer permissions, a publish decision, and a recovery path.

Eight production-readiness checks for one AI workflow
Readiness area Question to answer Example evidence to collect
1. Outcome What useful result must the workflow deliver, and for whom? Defined review and publication outcome; approval or rejection recorded
2. Data and system boundary What data enters each component, who can access it, and what can leave the system? Source page, glossary, locale, provider access, storage, and retention map
3. Evaluation How will quality be tested repeatedly rather than judged from memory? Representative pages, error rubric, glossary checks, reviewer notes, and regression cases
4. Security and permissions Can the workflow propose or perform an action beyond the user's authority? Role-based review and publish permissions, least-privilege credentials, and audit record
5. Reliability What happens on timeout, invalid output, duplicate processing, replay, or retried delivery? Schema validation, retry limits, idempotency, fallback, and recovery procedure
6. Cost and latency What cost, response-time, and throughput budgets are acceptable? Model cost per approved page, review time, queue depth, end-to-end latency, provider quota, sustainable throughput, and retry volume
7. Observability and ownership Can an operator understand what happened without asking the original developer? Logs, model and prompt versions, alerts, escalation owner, and support procedure
8. Rollout and rollback How will exposure increase, and which conditions stop or reverse the release? Internal use, limited pages or locales, monitored expansion, and versioned rollback

The table is a diagnostic prompt, not a certification checklist. A scoped workflow may need more or less evidence.

Eight production-readiness checks

1. Define the outcome and acceptance criteria

"Use AI to localize the site" is not an acceptance criterion. Name the user, the action, the allowed data, and the decision that follows. For this example, a reviewer approves or rejects a generated page before publication, and the system records that decision.

2. Map the data and system boundary

List the source system, model call, storage, reviewer interface, delivery host, and owner for each handoff. Record which data may be sent to the approved model provider, where it is retained, and who can access it. The NIST Generative AI Profile is a useful lifecycle reference for this context and risk work.

3. Build a representative evaluation set

An evaluation is a repeatable way to compare an output with an expected result or rubric. Start with representative and difficult inputs, define what counts as wrong, and keep production failures as regression cases. OpenAI's evaluation best practices describe one provider-specific approach; the method should remain understandable if the model or vendor changes.

4. Restrict permissions and side effects

Security and output validation are separate gates. Decide which data may leave the system, who can invoke the workflow, which actions the model may propose versus perform, and whether tool calls inherit the user's permissions. Validate structured output against a schema, sanitize or escape generated HTML for its rendering context, and authorize every side effect independently of the model output. Tools should run with scoped, least-privilege credentials rather than unrestricted user or application access. The OWASP LLM Top 10 is a risk taxonomy, not a full security audit.

5. Design failure and recovery paths

Plan for timeouts, rate limits, malformed output, duplicate processing, replay, retried delivery, and partial publication. Validate the schema before a side effect, cap retries, make retries idempotent, and give a person a clear recovery path. In the localization example, a failed page should stay unpublished while the reviewer can retry or restore the last approved version.

6. Set cost, latency, and capacity budgets

Cost, latency, and capacity are production constraints, not proposal footnotes. Track model cost per approved or published page, human review time, queue depth, end-to-end processing time, provider quotas, sustainable throughput, publish failure rate, retry volume, and the cost of duplicate work. A workflow that is affordable and fast for ten pages may still fail at ten thousand pages or when human review becomes the bottleneck. The acceptable numbers depend on the workflow's value and risk; agree the budget before expanding exposure.

7. Add observability and operational ownership

Log the input and output identifiers, model and prompt versions, validation result, reviewer decision, publish action, and recovery event. Define the operator, escalation owner, alert threshold, and support procedure so someone can understand what happened without asking the original developer. This is where who should own the deployment becomes an operational decision, not a job-title debate.

8. Roll out gradually with stop conditions

Write down which quality, latency, cost, or failure signals pause expansion. A rollout is more controlled and reversible when its exposure and stop conditions are explicit; a named owner must accept the residual risk before expansion.

  1. Offline evaluation. Run representative and difficult cases without affecting users.
  2. Shadow mode. Process real inputs without publishing or triggering external actions.
  3. Internal release. Let only the team review and approve results.
  4. Limited production release. Enable one locale, workflow, customer group, or content type.
  5. Measured expansion. Increase exposure only while agreed quality, latency, cost, and failure metrics stay within budget.
  6. Rollback or pause. Stop on a defined failure and restore the previous version.

Worked example: a website-localization workflow

The useful path is not "prompt in, page out." It is a chain of decisions and evidence:

  1. Source

    Capture content, source version, and route.

  2. Context

    Attach locale, glossary, and source metadata.

  3. Authorize

    Enforce reviewer and publisher permissions in application code.

  4. Generate

    Produce a versioned draft.

  5. Validate

    Validate structure, required terminology, and publication rules.

  6. Review

    Approve or reject with notes.

  7. Publish

    Publish the approved version, monitor it, and retain a rollback path.

Illustrative sequence: a production workflow keeps review, publication, and operational recovery attached to the model call.

Useful measures include cost per approved or published page, percentage approved without substantive edits, review time per page, queue depth, end-to-end latency, provider quota, sustainable throughput, retry volume, and publish failure rate. Two practical failure cases are a required glossary term translated incorrectly, causing reviewer rejection, and a timeout after validation but before publication; both should be visible, recoverable, and included in later regression tests.

Illustrative release boundaries for one localization workflow
Decision Illustrative rule Evidence
Scope Begin with one locale and selected page types. Page and route inventory
Publication gate Only a reviewer-approved version may go live. Approval record and version identifier
Quality gate Required fields, glossary terminology, and route checks must pass. Validation and reviewer results
Failure rule Invalid output or a timeout cannot replace the currently approved page. Job log and recovery test
Expansion Add pages or locales only while quality, latency, and cost stay within budget. Rollout review
Stop condition Pause on an unapproved publication, unrecoverable route conflict, or sustained budget breach. Alert and incident record

When to continue, narrow, redesign, or stop

These four outcomes keep a readiness review from collapsing into "ship" or "do not ship."

Four possible readiness decisions
Decision Use it when Next action
Continue Critical checks remain inside the agreed quality, risk, and operating budgets. Expand by one rollout stage.
Narrow One locale, page type, user group, or action is ready, but the broader scope is not. Reduce the deployment boundary.
Redesign Failures, permissions, evidence, or ownership cannot be isolated clearly. Change the workflow or architecture and retest.
Stop Expected value no longer justifies the risk, cost, or operating burden. Close the pilot and document the findings.

Frequently asked questions

What makes an AI prototype production-ready?

The scoped workflow meets its agreed outcome, risk tolerance, and operating constraints, with evidence that someone else can run and support it.

Can a system be production-ready if every output needs human review?

Yes. Human review is compatible with production when the rubric, permissions, logging, escalation path, and reviewer capacity are explicit. The problem is ad hoc judgment by the original builder.

How many examples should an AI evaluation set contain?

There is no universal number. Start with representative, difficult, and known-failure cases; add production failures as regression cases; and expand until the set covers the decisions and risks in scope.

What should be monitored after deployment?

Monitor quality signals, latency, cost, retries, failure rates, reviewer decisions, permission events, and rollback triggers. The right thresholds depend on the workflow's value and risk.

When should an AI pilot be stopped rather than expanded?

Pause when agreed quality, cost, latency, risk, or recovery thresholds are breached and the team cannot restore them inside the current boundary.

Sources and further reading

All articles