# Decision — Fake model output must never silently reach a real run (fail closed)

> **Spike:** [#1830](https://dev.azure.com/Aeonic-Tech/Manuscript%20Translator/_workitems/edit/1830)
> — "confirm whether fake OCR/translation output can silently reach a real (non-dev) run".
> **Parent feature:** #1365 (Core Translation & Review Experience).
> **Outcome:** a decision (expected vs. gap) + a scoped follow-up. This record promotes the spike note
> to `docs/design/` because the decision **changes a documented invariant** — the runbook's current
> "RunPod not configured → Fakes (dev/test default)" default (`docs/runbooks/runpod-model-processing.md`
> §TL;DR) — for non-development environments.
>
> ✅ **Implemented in [#1832](https://dev.azure.com/Aeonic-Tech/Manuscript%20Translator/_workitems/edit/1832)**
> (successor to #1830). `AddBookTranslation` took an `allowFakeModelBackend` gate and **threw** when the
> fakes branch was reached with the gate `false`; the API + Worker hosts derived the gate from
> `IHostEnvironment`. *(This gate was superseded by #1840 below.)*
>
> ✅ **Superseded / hardened by [#1840](https://dev.azure.com/Aeonic-Tech/Manuscript%20Translator/_workitems/edit/1840):
> the fakes were removed from the product entirely.** There is now **no fake model backend anywhere in
> `Infrastructure`/`Api`/`Worker`** — so fake output cannot be served *by construction*, which is a
> stronger guarantee than the #1832 startup gate (now removed along with `allowFakeModelBackend`).
> `AddBookTranslation` registers real model stages only when RunPod (or a local OCR worker) is configured
> and **nothing** otherwise; the host still starts (the `JobProcessor` resolves the pipeline lazily) so a
> translation attempt fails loudly rather than crashing boot. Deterministic model-stage **test doubles**
> now live in the **test assembly** (`BookTranslation.Tests/TestDoubles`) and are injected into the real
> host via `WebApplicationFactory.ConfigureTestServices` (`TestModelBackend.AddModelBackendDoubles`) — the
> product ships zero fakes and CI still runs fast + offline. See **§Removal (#1840)** below.

## TL;DR — decisions

| Question | Answer | Evidence |
|---|---|---|
| Which backend served the 2026-07-04 "Test" screenshot (`(fake-translation) [to:en] …`, OCR `94%`)? | **Fakes** (Milestone-1 placeholders). | The `(fake-translation) [to:…]` prefix is emitted **only** by `FakeTranslationService` (`FakeAdapters.cs:204`); the `94%` is the hardcoded `PageConfidence: 0.94m` in `FakeOcrService` (`FakeAdapters.cs:140`). Both are registered **only** in the final `else` branch of `DependencyInjection.cs:172–181`. |
| Was that a bug, or expected? | **Expected M1 dev-default.** No `BookTranslation:RunPod:*` was configured for the host that served it, so the precedence chain fell to fakes by design. | API host sets `runPod = null` unless **both** `BookTranslation:RunPod:ApiKey` **and** `:EndpointId` are set (`Program.cs:81–99`). M1 is "fake adapters" by design (CLAUDE.md status). |
| Is there a misconfiguration gap **today**? | **No.** There is no deployed prod/UAT environment that is *expected* to route to RunPod but silently falls back — RunPod is opt-in and stood up on demand (`scripts/runpod-up.ps1`). No environment silently regressed → **no Bug filed.** | RunPod bring-up is a manual, on-demand script; no always-on real environment exists yet. |
| Is the current signal sufficient? | **No.** The only signal is one startup log line (a `Warning`), watched by nobody; `ModelBackendInfo.IsReal` is surfaced on **no** API endpoint and **nowhere** in the UI. The lone in-band tell is the literal `(fake-translation)` prefix — easy to mistake for real content. | Whole-repo grep for `ModelBackendInfo\|IsReal\|ModelBackend` returns only 3 backend files (DI, announcer, runbook) — zero API/UI surface. |
| Fix direction (per product owner)? | **Remove the silent fallback — fail closed.** A non-Development host must be explicitly configured with a real backend and **refuse to start** if it isn't, rather than silently registering fakes. Fakes stay reachable **only** by explicit opt-in (Development/Testing, or an explicit flag). | Scoped as follow-up **PBI (successor to #1830, child of #1365)**; not built in this spike. |

## The question

A Review-page screenshot (project "Test", page 15/30, 2026-07-04) showed the translated pane reading
`(fake-translation) [to:en] Lorem ipsum dolor sit amet, pagina 15, pars 2.` with OCR confidence pinned
at exactly `94%`, while the project read "Translated" / "Complete". Both strings are Milestone-1 stub
hallmarks. Is this the expected dev default, or a wiring gap that could let a real, paying run silently
fall back to fakes with no user-visible signal?

## Current state (read from the repo, not assumed)

**Backend selection** — `AddBookTranslation(...)` picks by precedence **RunPod > local Python OCR
worker > fakes** (`src/BookTranslation.Infrastructure/DependencyInjection.cs:114–181`):

- `runPod is not null` → all model stages on RunPod (`:121–157`).
- else `ocrWorker is not null` → local OCR worker + **fake** language/translation (`:158–171`).
- else → **all fakes** (`:172–181`), `ModelBackendInfo(ModelBackend.Fakes, "FAKES …")`.

**How the API decides** — `runPod` is non-null only when **both** `BookTranslation:RunPod:ApiKey` and
`BookTranslation:RunPod:EndpointId` are present (`src/BookTranslation.Api/Program.cs:81–99`). Absent
either, the host silently constructs the fakes graph. There is no environment check — a **Production**
host with a missing/typo'd key degrades to fakes exactly as a dev box does.

**The signal** — `ModelBackendAnnouncer` (`ModelBackendAnnouncer.cs`) logs once at startup: `Information`
for RunPod, `Warning` for fakes/local-worker. `ModelBackendInfo.IsReal` (`true` only for RunPod) exists
but is consumed **only** by that logger. It is on no endpoint and in no screen.

## Analysis

The screenshot is **benign and expected**: a dev run with no RunPod config, correctly served by the M1
fakes. Nothing regressed; no Bug.

The **latent risk** is real and forward-looking. The moment M2/M3 stand up an always-on paying
environment, the *same* silent-degradation path means a missing RunPod key there would serve
`(fake-translation)` placeholder text to a paying user — and, once real billing meters compute (M1
fakes meter `$0`, so today the wallet isn't charged — see `docs/design/cost-estimate-accuracy.md`),
risk charging for placeholder output. The only thing standing between that and a user is a startup
`Warning` nobody reads and a parenthetical prefix easy to miss.

Two ways to close it were considered:

1. **Surface the active backend** (expose `ModelBackendInfo`/`IsReal` on the API + a "placeholder
   output — not a real translation" banner in Review). Makes the state visible but still *permits* a
   real environment to run fakes — it only labels it.
2. **Fail closed** (chosen). Make fake selection impossible by omission in a real environment: a
   non-`Development` host that isn't explicitly configured with a real backend **throws at startup**
   rather than registering fakes. Visibility becomes moot because the unsafe state cannot occur in
   prod.

Option 2 is the product owner's directive ("remove all fallback behavior") and the stronger guarantee.
A visible indicator (option 1) is a reasonable *complement* for dev/demo clarity but is secondary.

## Decision

**Remove the silent fake-backend fallback; fail closed.** Concretely, for the follow-up PBI:

- In a **non-Development** environment (`!IHostEnvironment.IsDevelopment()`), if no real model backend
  is configured (no RunPod, no real OCR worker), the host **must throw at composition** with an
  actionable message naming the missing `BookTranslation:RunPod:*` keys — it must **not** register the
  fakes graph.
- Fakes remain selectable **only** by explicit opt-in — Development/Testing environments, or an explicit
  `BookTranslation:AllowFakeModelBackend=true` (or equivalent) flag — so tests and the bare `dotnet run`
  dev loop are unaffected, but fakes can never be reached by config *omission* in a real deployment.
- **Do not delete `FakeAdapters`.** The M1 test suite and the dev loop depend on them; this is about
  making their selection explicit and impossible-by-accident, not removing the code.
- Update `docs/runbooks/runpod-model-processing.md` (the "Fakes = default when RunPod not configured"
  TL;DR row) and CLAUDE.md to state the fail-closed invariant.

A user-visible backend indicator (surfacing `IsReal` on the API/UI) is **optional and secondary** — it
may be folded into the same PBI as a dev-clarity nicety or split out, at the implementer's discretion.

## Verification (evidence, not assumed)

- `FakeAdapters.cs:140` — `OcrResult(PageConfidence: 0.94m, …)`; `:204` — `$"(fake-translation) {into} {sourceText}"`. The screenshot's two hallmarks are these exact literals. ✔
- `DependencyInjection.cs:172–181` — fakes registered only in the no-RunPod/no-worker `else`. ✔
- `Program.cs:81–99` — API leaves `runPod` null (→ fakes) unless both RunPod key + endpoint are set;
  no environment gate. ✔
- Whole-repo grep `ModelBackendInfo|IsReal|ModelBackend` → 3 files, all backend; **zero** API/UI
  consumers. ✔
- `docs/runbooks/runpod-model-processing.md` §TL;DR documents today's "Fakes = dev/test default,
  WARNING at startup" invariant that this decision revises for non-dev. ✔

## Implementation (#1832)

The follow-up shipped exactly the fail-closed decision above. The shape:

- **`AddBookTranslation(..., bool allowFakeModelBackend = true)`** — the composition root gains a gate.
  Its fakes `else` branch now throws an `InvalidOperationException` *before* registering any fake adapter
  when the gate is `false`, with an actionable message naming `BookTranslation:RunPod:ApiKey` /
  `:EndpointId` and the `BookTranslation:AllowFakeModelBackend` opt-in.
  (`src/BookTranslation.Infrastructure/DependencyInjection.cs`.)
- **The hosts derive the gate from their environment**, so fakes are opt-in only:
  ```
  allowFakeModelBackend =
      Environment.IsDevelopment()
      || Environment.IsEnvironment("Testing")
      || config["BookTranslation:AllowFakeModelBackend"] == true
  ```
  (`src/BookTranslation.Api/Program.cs`, `src/BookTranslation.Worker/Program.cs`.)
- **Why the library default is `true`, not `false`:** `AddBookTranslation` is an `Infrastructure`
  extension called both by the product's hosts (which always pass the environment-derived gate) and by
  bare in-process test callers (which have no `IHostEnvironment` and want the fakes). The safety lives at
  the host seam — there is no way to run the shipped product without going through the API or Worker host,
  which is what makes the invariant impossible by construction — so a `false` default would only churn
  tests without adding product safety.
- **`FakeAdapters` are untouched** — only their *selection* is now explicit.
- **Verified by `ModelBackendFailClosedTests`** (`src/BookTranslation.Tests`) + the runbook/CLAUDE.md
  updates; the existing suite stays green because it runs under `Testing`/bare in-process.

## Removal (#1840) — no fake model backend at all

Product direction (2026-07-04): *"remove any remnant of the fakes — let's not have the fakes at all."* The
#1832 startup gate made a real deploy *refuse to serve* fakes; #1840 goes further and deletes the fakes, so
there is nothing to serve. The shape:

- **Deleted** `Infrastructure/Fakes/FakeAdapters.cs` (fake normalizer/preprocessor/layout/OCR/language/
  translation) and removed the fakes DI branch + the `allowFakeModelBackend` gate + `ModelBackend.Fakes`.
  Grep for `Fake` across `Infrastructure`/`Api`/`Worker` model code returns zero hits (the remaining
  `FakeCardPaymentGateway`/`FakeFxRateProvider` are unrelated, config-gated *payment/FX* test gateways).
- **`AddBookTranslation`** now registers real model stages only when RunPod (or a local OCR worker) is
  configured, and **nothing** otherwise (`ModelBackend.None`). The one genuine no-op that remained — page
  preprocessing — is now real product code, `PassThroughPagePreprocessor` (not a "fake"). The normalizer is
  always the real `PdfPageNormalizer`.
- **`JobProcessor` resolves the pipeline lazily** (from `IServiceProvider`, per job) instead of taking it in
  its constructor, so a host with no backend still *starts* — a job then fails with a clear error rather
  than crashing host startup. (Graceful "GPU starting" UI is #1833.)
  - **`ITranslationPipeline` registration itself is deferred, or lazy resolution is not enough (Bug #1853).**
    In **Development**, `WebApplicationBuilder` turns on `ServiceProviderOptions.ValidateOnBuild`, which walks
    every singleton's constructor **call site at host-build time**. A plain `AddSingleton<ITranslationPipeline,
    TranslationPipeline>()` therefore had its call site validated eagerly and threw on the unregistered
    `ILayoutDetector` — crashing boot *before* the lazy `JobProcessor` path ran, so the "host still starts"
    guarantee held only in Production (where `ValidateOnBuild` is off). Fix: in the **no-backend** branch
    `AddBookTranslation` registers `ITranslationPipeline` via a **factory that throws an actionable "no model
    backend configured" error on first use**. A factory call site is opaque to `ValidateOnBuild`, so the host
    boots; the first resolution (JobProcessor per job, or a `/process*` endpoint per request) fails loudly.
    When a backend *is* configured, the plain type registration is kept (unchanged working path). The test
    host always hits the no-backend path, so `TestModelBackend.AddModelBackendDoubles` re-registers the real
    `TranslationPipeline` over that factory once its stage doubles are in place.
- **Tests** keep running fast + offline: the deterministic model-stage doubles moved to the **test
  assembly** (`BookTranslation.Tests/TestDoubles/FakeModelAdapters.cs`). Direct-construction tests
  (`PipelineTestHarness`, `ResumabilityTests`) reference them there; host-based integration tests inject
  them via `WebApplicationFactory.ConfigureTestServices` (`TestModelBackend.AddModelBackendDoubles`). The
  product contains **zero** fake model code. Full suite stays green (544 tests, 0 skipped).

## Follow-up

- ✅ **#1833 (shipped):** model-backend readiness is surfaced on the API and in the SPA — the graceful
  counterpart to "the host starts even without a live backend." Shape:
  - **`GET /v1/model-status`** (anonymous) reports `{ backend, isReal, ready, state, detail }` where
    `state ∈ ready | starting | unconfigured` (`docs/api-contract.md` §5). Backed by an always-registered
    `IModelBackendReadiness` that maps the configured `ModelBackendInfo` + a live health probe to a status.
  - The health probe is `IModelHealthProbe`, **implemented by `RunPodClient`** via the worker's `{op:"health"}`
    op on the *same* `/runsync` transport that serves real OCR/translation — so it is **pod/serverless-agnostic**
    (works unchanged on the dev persistent pod and a production serverless endpoint; only `BaseUrl`/`EndpointId`
    differ). The probe result is TTL-cached + coalesced so a polled endpoint triggers at most one backend
    health call per window (bounding serverless worker wake-ups).
  - The SPA polls it app-wide (`ModelReadinessProvider`), shows a top **"GPU infrastructure is starting, please
    stand by…"** banner (`ModelReadinessBanner`) while not ready, and **disables the process action** until the
    backend reports ready. No new persistent config toggle — behaviour is automatic.
  - Real-run (pod stopped → banner + gate; pod started → clears) is the same code path as the tested
    ready/starting probe outcomes; exercising it against a live pod is billable RunPod spend, deferred to an
    on-demand bring-up.
