← All drafts
The roadmap for taking The Offer from the screenplay to a finished AI-generated
short. Working order is top to bottom; checked boxes are done. Conventions live
in `AGENTS.md` and still apply to every step here.
1. The number is the reveal. The 400-gulden figure appears on screen exactly
once: the INSERT of the unfolded contract in Katharina's hands, near the end
of the courtyard scene (and the burn that follows it). In every other shot,
in every scene, the contract's writing is angled away, back-to-camera, or
folded. Reject any generation that leaks a legible sum earlier.
2. All in-frame text is period German. Blackletter (Fraktur/Schwabacher) for
print, a 16th-century manuscript hand for the letter, ledger, and contract
(the reveal reads `400 Gulden jährlich`). Reject any generation showing
English words anywhere.
Higgsfield is the production platform for everything generated — picture,
voice, music, sound design. The CLI is installed and authenticated (`plus`
plan, 1,200 credits granted monthly); companion skills live in
`.agents/skills/`. Model inventory, measured prices, and what this costs us
are in Services below.
One consequence is load-bearing and is recorded here so it is not discovered
late: no engine reachable through Higgsfield accepts phoneme (IPA) input.
Period pronunciation is therefore forced by respelling the recording script,
not by a pronunciation dictionary — see Phase 2a. That is a real capability
we are choosing to give up in exchange for a single platform.
- [x] Current draft: `the-offer.fountain` (10 September 2026) — runner-intercut morning, evening council, study refusal, courtyard reveal. Two revisions since the 30 July draft. The 7 Aug revision changed the morning only and moved no dialogue: the market beat cut and the alley moved into its slot, so the runner crosses town in three exteriors instead of four (Mabel's direction — fewer city shots; reasoning in `shots-morning.md`). The 10 Sep pass rewrote the evening council heavily — nineteen changes across two sittings, plus a register pass — and every one of them is carried into `the-offer-de.fountain` and `shots-evening.md`; the write-up is in that sheet's Read-through pass and Mabel's own pass sections. It also emptied most of the English title page's Notes block, whose argument had drifted into duplicating `resources/pronunciation-guide.md` §3.13, and discharged the German header's `gekratzet` transcription flag, which `resources/table-talk-4690.md` had already closed.
- [x] PDF render working: `~/.local/bin/screenplain the-offer.fountain the-offer.pdf` (11 script pages ≈ 11 min).
- [ ] Final read-through pass for dialogue that doesn't land (the Lufft "not a purse" line was one such fix — and it was cut outright on 10 Sep 2026, so it is no longer in the script to point at; read for others). The evening scene got such a pass on 10 Sep 2026 — six changes at Mabel's direction, all propagated to both fountains, `shots-evening.md` and the timings, and all written up in that sheet's Read-through pass section: (1) Lufft's "Not a purse. A contract…" cut, so the terms are not explained before Luther hears them; (2) "selling what belongs to Wittenberg" for "what Wittenberg was given free"; (3) Lufft turns to the Old Printer on "your wife," which E-17's pan now has to serve; (4) the sum reactions carry its size rather than the printers' poverty, measured against the Elector's stipend; (5) "His words" for "His ink"; (6) "It is decided" for *"Are we agreed?"* — Lufft does not ask. Two knock-ons left open in the sheet: the Old Printer's "bend him with a purse" now goes unanswered, and the word contract no longer occurs in the evening. A second pass the same day, typed by Mabel straight into the rendered screenplay and folded back in, made thirteen more changes to the same scene — nine of her own rewrites and four lines she left to be written; two of them collided with the language-seam note in `resources/pronunciation-guide.md` §3.13 and with the scene's own chronology, and both were fixed rather than left standing. All of it is written up in `shots-evening.md`. The morning, study and courtyard have not had the pass.
- [ ] Decide the title sequence before lock — where the film's own title falls, in which language, and whether it is a card or an end title (the case is under Phase 6). It is a shot like any other and belongs in the fountain, not bolted on in the edit. *(7 Aug 2026: language and name are settled — English, and The Offer stands. Placement is narrowed to the cut out of M-15 but not picked. Full state under Phase 6.)*
- [ ] Declare the script locked. After lock, any script change must propagate to the generation sheets before generating from them.
Two deliverables, in a fixed order: 1a — `the-offer-de.fountain`, every
spoken line in Early New High German (decision of 1 Aug 2026; English
subtitles are authored later in the edit, and the English fountain stays the
story's source of truth). 1b — one generation sheet per scene, quoting
its dialogue verbatim from the German script. The order is fixed because the
sheets quote the German word for word: 1a must pass verification before the
remaining sheets are written or re-quoted.
- [x] Write `the-offer-de.fountain` — every spoken line in ENHG: 16th-century grammar, lexis, and forms of address (Ihr / Herr Doctor), spelling lightly normalized, the English line kept beside each German one. *(Drafted 1 Aug 2026 against the unlocked 30 July draft — reconcile at lock.)*
- [x] Luther's refusal speeches from the source, not from us — adapted from the Tischreden German, not back-translated. *(WA TR 4, Nr. 4690; salary parallel Nr. 5151. Citations inline at each speech.)*
- [ ] Verification pass by a native or period-literate reader for accidental modernisms and false archaisms. Nothing is recorded, and no sheet quotes a line, until this passes. *(Scene-scoped exception, 1 Aug 2026: the morning scene's six lines were checked and `shots-morning.md` re-quoted in German — the morning proceeds ahead of this gate and of script lock; the pass stays open for the other scenes.)*
Format per `shots-morning.md`: intro note naming the draft it was written
against, a BOILERPLATE block, a shot grid (one row = one generation), a
continuity checklist. The Audio column quotes the German script verbatim —
no shortened speeches. The study and courtyard sheets also need the
scratch-track timings (Phase 2a), so the practical order is: 1a → evening
re-quote → scratch track → study and courtyard.
- [x] `shots-morning.md` — 11 shots, M-01..M-15 less the cut M-06, M-10 and M-13 and the retired M-14 (runner intercut; Audio re-quoted from the German script 1 Aug 2026). IDs are stable names, not cut order: M-09 plays before M-07. M-10 went 5 Sep 2026 at Mabel's direction and took the lampblack exchange with it, so the morning is four scripted speeches, not six. M-13 went 7 Sep 2026 ("delete m13, we dont need it") — the shop going quiet, never rolled, no dialogue of its own, and its beat had already died with the instruction that nobody is on the press. **M-14 folded into M-12 the same day* ("we can also combine m14 with m12"*), so the door and the letter are one shot; its German VO survives the merge and the speech count is unchanged. Reasoning for both in `morning-sheet-history.md`.
- [x] M-12's camera — DECIDED 9 Sep 2026: the static wide, trimmed. Mabel picked `M-12-53`, which is `M-12-47` (the wide, no shoulder) with the first second cut off the head — 241 frames to 217, 9.042s — so the answer was neither rolled family as it stood but a local trim of one of them, the same shape as `M-11-12`. It is keyed at `approved/M-12/M-12-53-f14847ac.mp4` and cut into `morning-scene-35`. The morning now has ten of eleven shots picked; only M-15 is unshot. The record of the decision is in `morning-sheet-history.md`; the M-12 row of `shots-morning.md` still reads undecided, and the choice is the shot and has not caught up. What follows is the state as it stood before the pick, kept because the two families and their costs are still the reasoning behind it. The combined row is 11s of screen time, the first to pass Veo's 8s ceiling. Not a tooling limit: `seedance_2_5` has rolled a 10s take and an 18s one, and M-15 has asked for 11s of Gen all along. **48 takes in, the choice is no longer between descriptions but between two families that exist on disk**, both locked off, both 10s, both free: static over-the-shoulder (`M-12-40`…`-46`, Lufft's back and near shoulder filling the lower-left foreground) and static wide with no shoulder (`-47`, `-48`, the same room from inside with the foreground figure removed). The stillest of each is `-45` at 5.2 drift and `-47` at 4.2 — the two stillest takes the family has made, so this is a framing choice and not a quality one. The internal-cut version her original brief describes still has a written prompt (`probes/video/m12-internal-cut-prompt.txt`) and has never been rolled; the over-the-shoulder push blocked in Blender on 7 Sep (`refs/blender-movies/printshop-shoulder-push-01.mp4`) is not among the rolled options — every take of the shoulder family is static. Until one is picked the morning has no credit estimate — the old ~631 for one take of everything is withdrawn rather than restated on a guess. Either way the insert grammar is a casualty — shallow focus is reserved for M-03 and the letter landing, and one continuous move from door to table reads the letter in the same breath as the room. The two families are not the same prompt minus a shoulder (measured 9 Sep 2026): the wide also drops six clauses the shoulder version keeps — the explicit does NOT pan, tilt, orbit, track, dolly, crane, zoom list, *NO deliberate camera movement, the account book's closed, front cover up, clasp fastened, the Apprentice's must not be confused with the brown-haired Runner*, the four-line NO ONE JERKS block, and the note that the Runner leaves frame because the camera is static. It still measured the lower drift. If the wide is picked, the next roll should carry those clauses back rather than inherit `-47`'s trimmed text.
- [ ] Roll M-15 — prompt written 9 Sep 2026, unrolled, **and now the only morning shot with nothing generated for it.** 11s at 720p on `seedance_2_5`: 71.5 credits. It was written as free, but Unlimited has charged every roll since 10 Sep (see § 2b, Unlimited stopped being free). One thing blocks it, and it is Mabel's: the approved `M-12-53` has the Old Printer beneath the drying cords with a printed sheet in both hands, and M-15's row has him at the counting desk, so the hard cut between them walks him across the room in no time. **M-12's approval closed one of the three ways out** — re-staging M-12 would now mean unpicking an approved take — so it is down to two: end the pan on him where M-12 left him and lose the counting desk, which is the thematic point of the shot, or accept the transit and let the pan imply the time. Full statement of the choice in the M-15 row of `shots-morning.md`.
- [x] `shots-evening.md` — E-01..E-19 (printers' council).
- [ ] Re-quote `shots-evening.md`'s Audio column from the German script (E-01..E-19 were written with the English dialogue). **Blocked on the § 1a gate** — the evening has no scene-scoped exception, unlike the morning.
- [ ] **Seven open questions on the evening scene, flagged for a later session to ask* (10 Sep 2026, Mabel: "put a flag on it and a dif session can ask later"*). Indexed under OPEN QUESTIONS at the top of `shots-evening.md`, and repeated on the rows they belong to. They include the § 1a gate above, whether E-19 plays in darkness now that the candle goes out in E-18, and two lines written by a session rather than by her that are candidates for cutting. None is an oversight; none should be settled without asking her. The sheet's list also records one problem she was shown and waved off — the 2× picture-to-dialogue shortfall — so it is not re-raised as though new.
- [ ] Write `shots-study.md` — the refusal, from the printers filing in through "Wolf! Ask my wife to come out to the courtyard." New elements: Luther (mid-fifties, heavyset, clean-shaven, black scholar's cap), the study set (books, proofs, the worn Bible on its desk), the contract prop carried over from E-16 and never unfolded. Luther's speeches are quoted verbatim from the script — no re-adaptation at sheet level.
- [ ] Write `shots-courtyard.md` — golden hour, wash kettle and fire, Katharina (about forty, linen coif, ring of keys at her belt). Contains the film's only two number-visible shots: the INSERT (`400 Gulden jährlich` buried mid-clause in a clerk's hand) and the burn (the number the last thing to burn); every earlier shot keeps the paper's face away from camera, including Luther's unfold-and-read (we stay on his face).
- [ ] Retire the parked sheets (`shots-scene1/2/3.md`) in a single commit whose message names their replacements.
- [ ] Add each new sheet's stem to `SCRIPT_ORDER` in `build.py`.
Both get decided by cheap probes, not marketing pages. Speech timings must
exist before picture is generated, because generated clips have fixed
durations. Concrete services, prices, and tradeoffs live in Services
below.
- [ ] Pronunciation is forced by respelling, not by IPA. Dialogue is TTS (AI-only production, 1 Aug 2026) and no one here speaks ENHG, so period forms have to be made to happen — stock modern-German TTS silently "corrects" them. No Higgsfield engine takes phoneme input, so the lever is the one thing we fully control: the spelling of the text we send. Write the word so the engine's own German reading of it lands on the period sound. This needs no vendor feature and works identically on every engine, which is why the platform decision does not endanger it.
- [ ] Two versions of every line, and the split is the whole method. `the-offer-de.fountain` stays readable ENHG — it is what the Phase 1a verifier checks, what the subtitles are written from, and the story record. A separate spoken script carries the respellings actually sent to the engine. Never conflate them; never let a respelling migrate back into the fountain. Generate the spoken script from the fountain so it can be rebuilt when a line changes.
- [ ] Choose the engine among Higgsfield's four (candidates and the ElevenLabs caveat: Services — Voices). Probe with a few test lines — one long Luther speech, one Lufft, one Katharina — plain first, then respelled on the failures. Judge: how much respelling per line does it take to hold the period forms, and does it act, not just read? Score word-level pronunciation fixes separately from phrasing and stress fixes; respelling can only reach the former, and the rest is delivery direction (`qwen_audio_tts` takes an `instruction` field for exactly that). Probe audio and any scratch harness stay out of git (`probes/` is gitignored).
- [ ] Decide how far the reconstruction goes, and write it down. Per `resources/pronunciation-library.md` there is no recording and no consensus reconstruction of 1539 East Central German — the sources support a defensible flavor, not a phoneme-by-phoneme rebuild. So pick a small number of features, apply them consistently across every character, and record the choice; anything beyond what the sources carry is a dramatic choice and gets marked as one per `AGENTS.md`.
- [ ] Scratch track first, regardless of route: rough TTS of all dialogue, time every speech, write the durations into the sheets' rows before the study and courtyard sheets are written. *(Harness `scratch_track.py` → `scratch-timings.md`; estimate pass done 1 Aug 2026 — 59 speeches, ~8.4 min of dialogue, longest speech ~29 s, so long speeches will split across shots. Regenerated 5 Sep 2026 when M-10's cut took the lampblack exchange out of the German script: 57 speeches, ~8.3 min, and the morning's MOR IDs renumber (see `morning-sheet-history.md`). Re-run for measured timings against the chosen Higgsfield engine; estimates don't satisfy this box.)*
- [x] Fix the morning's over-long shots — found 3 Aug 2026 by checking `scratch-timings.md` against `shots-morning.md`: M-10 allots 5 s for 8.8 s of dialogue, M-14 4 s for 5.4 s, M-15 5 s for 5.9 s. M-04 fits. Estimates, so M-15 is inside error, but M-10 and M-14 are not close calls. *(Fixed in the sheet the same day: M-10 5→9 s, M-14 4→6 s, M-15 5→8 s — the sheet's header records it. Re-check against measured timings when they exist.) *M-10 was cut on 5 Sep 2026 and its dialogue with it**, so only the M-14 and M-15 fixes are still load-bearing — and the flexible-duration model M-10's 9 s forced is no longer required by anything in the morning. M-14 itself was folded into M-12 on 7 Sep 2026 and M-13 was cut the same day; M-14's dialogue survives the merge, so the 6 s that fix bought is still doing its work, just inside M-12's 11 s. The flexible-duration problem is back, and worse: 11 s is past what one generation gives on either Seedance model, which is the open question on M-12's Gen cell.
- [ ] Audio meets picture as VO/dub (the default stands): framings favor listeners, hands, and profiles so lip-sync is rarely needed; native audio generation is rejected — no video model can be directed to speak ENHG. Count the unavoidable on-camera lines when the study sheet exists, then decide on lip-sync tooling.
- [ ] One pronunciation guide, written once with the Phase 1a verifier and applied to every character — the ensemble must sound like one century.
The service is settled (the platform, above); what is open is which
Higgsfield model does the bulk and which does the hero shots.
- [ ] Hard requirements — still worth checking, now per model rather than per vendor: (1) multi-reference image-to-video conditioned on our approved stills, not single-frame conditioning — this is the one that reordered the picks, so check the exposed parameters, not the marketing; (2) clip length ≥ the longest speech-carrying shot per the scratch-track timings — Veo is 4/6/8 s only, so long speeches get split across shots in the sheets until a 30 s-class model is reachable; (3) no forced watermark, commercial use permitted — **confirm once for Higgsfield and record it here**, since the terms now come from one vendor, not four; (4) will generate period religious/historical subjects — probe with a Luther-likeness prompt.
- [ ] Budget envelope: ~70 shots × ~6 s × 3 takes ≈ 1,200–1,500 generated seconds. In Higgsfield credits a full pass ≈ 1,400 on Veo 3.1 Lite, 2,100 on Kling 3.0 sound-off, 3,850 on Veo 3.1, 6,300 on Seedance 2.0 — so no full pass fits inside one month's 1,200-credit grant, and passes must be either scheduled across months or paid up. *(Prior dollar framing, kept for reference: all-Kling ≈ $125/pass, mixed ≈ $300–400, all-Veo ≈ $600.)* Cost is not the deciding factor as of 3 Aug 2026 — the subscription can be raised, so the bake-off is graded on quality alone.
- [ ] Bake-off probe: the same three shots on `seedance_2_0`, `minimax_h3`, `kling3_0` and `veo3_1` — candlelit council interior, golden-hour courtyard, Luther close-up conditioned on his approved still. Grade consistency across a second shot of the same character first, then likeness, period look, light, motion, policy friction; results in `probes/video/` (gitignored). Pick a primary and a fallback — takes that keep failing on the primary get retried on the fallback.
- [ ] The pick: `seedance_2_5` for all video (Mabel's direction, 7 Aug 2026), with `seedance_2_0` as the fallback — the Seedance family is the only one exposing multi-reference conditioning, full case and the parameter table in Services. Then `minimax_h3`, `veo3_1` for the hero tier, `veo3_1_lite` for previz, `kling3_0` where a shot needs a length outside Veo's 4/6/8 s. All are the same CLI, so switching is a flag, not a rebuild — and the bake-off is one sitting, not four signups.
- [x] 2.5's reference gate — cleared by job references, 12 Aug 2026 (see Services). 2.5 rejects uploaded media with `IP check not finished`, and references are how likeness reaches a take; pointing at the job that made the picture uploads nothing and is never screened. `gen.py --ref <still>` does the translation, `#N` picks a panel of a turnaround sheet, and it refuses rather than falling back to an upload. **No take needs to fall back to 2.0 for reference reasons** — a still that cannot be referenced is a still to regenerate on the platform, not a reason to change model. Pass `--model` explicitly every time; the installed skills default to the dearest option in each category. An uploaded `--start-image` cleared the check, 19 Aug 2026. M-03's practical plate matches no job by construction — pixels no model made are the whole point of it — so it can never be an `image_references` entry, and `gen.py` refuses it rather than waiting four minutes for `IP check not finished`. Sent as a start image the same file went straight through, and the take begins on exactly those pixels. Treat it as one data point against the gate and not a repeal: a start image is not a reference, and it wants re-testing before it is relied on. What fixed staging on M-03 was neither prose nor a plate, 25 Aug 2026. The shot refused the same instruction twice in words — eight restatements of the heap is at the press, then 21,233 characters with six named failure conditions — and returned a heap across the room both times. What finally landed it was M-03-11: a free web roll with no plate of any kind, and the two things that made the difference are neither of them wording:
- The reference set was not what the thread believed. Match references by job id against `refs/provenance.json`; the composer's thumbnails are 48 px and the DOM carries no names. M-03-10 went out with five references, three of them character sheets, and two of those men duly turned up in frame holding one corner of the sheet each. Three sheets meant three candidates.
- The prompt on the wire was not the prompt in the box. The composer's prompt box is a rich-text editor, not a textarea: typing into it leaves the page holding its previous copy and submitting that, which is how M-03-10 went out with the 21,233-character prompt while the box showed 1,600. Set it with a real paste, then read the job back off the API and check `params.prompt` and `params.medias`. **Verify on the wire, never on the screen.**
So the reference set fixes casting, and getting the prompt onto the wire
is what lets the prompt fix anything at all. Neither costs a credit.
Plates are the untested lever, and they are aimed at what is left.
`m03-startframe-01` and `m03-endframe-01` are local composites of approved
stills (the arms and sheet of `wet-sheet-04` over `shop-day-26`'s near press;
M-03-09's landing frame with the press composited in behind it). They cost
nothing to build and have never been rolled. The one defect prose and
casting have not touched is in M-03-11 still: the sheet lands rotated about
90° — both headings along the top edge while he holds it, along the bottom
edge reading sideways once it lands. Foreshortening cannot move which edge a
heading sits on. That, and not staging, is what a pinned end frame is for.
Frame pinning is API-only, and pinning is not what costs, 25 Aug 2026.
Two facts measured on M-03 today, which together mean free and
plate-pinned do not intersect:
- The `seedance_2_5` **web composer has no first- or last-frame control at all**. The panel offers References, Extend Video, Prompt, Model, duration, aspect, Bitrate and Unlimited mode, and nothing else. Extend Video continues a video, it does not take a plate, and Motion Control is a different product (a Kling 3.0 feature, 7 cr on its own button, motion copied off a video reference; the model id behind it unconfirmed).
- The API price is unchanged by pinning: a 4 s 720p roll prices at 26 credits with no plates, with an end frame, or with both. The API is what costs, because Unlimited is web-only and has no param on the model.
So the trade is never "plates cost extra." It is one free web roll the
model stages for itself, against one 26-credit roll whose first and last
frames we decide.
Trap: the Unlimited toggle resets when you leave the Create Video tab
(25 Aug 2026). Stepping into another product and back kept the prompt and
all three references but silently turned Unlimited off — the button went
from Unlimited to 26. Re-read the button immediately before Generate and
confirm it carries no credit figure. This is the likeliest way a roll
believed to be free is charged.
Unlimited stopped being free on 10 Sep 2026 (found 18 Sep from
`higgsfield account transactions`). The ledger shows 254 `seedance_2_5`
rolls charged 0 between 19 Aug and 8 Sep. **Every roll from 10 Sep on, 33 of
them, was charged at 6.5 credits a second, 2,164 credits in all.** Most of
those were the evening, 1,885 net of refunds. The rolls were set up as
Unlimited, and the ledger shows no plan change, grant or reset between 8 and
10 Sep that would explain it. A session flagged this on 16 Sep (thread t21),
but the flag never reached this file, and the rolls went on being proposed as
free. Treat every video roll as paid at seconds × 6.5 credits. The
evening's 1,885 credits bought seven shots, about 270 a shot at roughly four
rolls each. The rest of the film at that rate is about 14,700 credits against
a balance of 3,876.5. Only a recent roll charged 0 in the ledger shows that
Unlimited works again. The trap above is not the explanation, because the
label was re-read before each roll and they were charged anyway.
- [x] The still pick: `seedream_v5_pro`, for every still (Mabel's direction, 12 Aug 2026) — character turnarounds, sets, props, and every regeneration or fix of one. It is the counterpart to the `seedance_2_5` video pick above: one model for stills, one for video, chosen once so a take is never quietly routed elsewhere. **Use Seedream and not other things* — this supersedes the older per-job model notes as routes*: `nano_banana_pro` for off-axis turnaround views (4 Aug) and for instruction edits of an approved still (4 Aug, qualified 7 Aug), and `flux_2` for cheap in-pose views. Those findings still describe real model behaviour and are kept as history below; they are no longer what to reach for. The reason this is the pick that matters is the gate directly above: a still generated as a **single Seedream job is referenceable whole by `seedance_2_5`**, while a `nano_banana_pro` edit or a local recut produces bytes that match no job — which is exactly what retired the 7 Aug Lufft and Young Printer sheets and cost two days of regeneration to undo. A fix that breaks the job match is not a fix. **Pass `--model seedream_v5_pro` explicitly**; never let a skill default choose.
Sequencing note, 3 Aug 2026. The morning slice of this phase depends on
nothing still open in Phases 0–2: character and set appearance do not move
when dialogue changes, the morning sheet is written and already re-quoted in
German, and every spoken line in that scene is VO or off-camera walla — so no
lip-sync, and none of the Phase 2a audio gates apply. (PLAN called the morning
"wordless" under Phase 4; "no sync dialogue" is the accurate reason it is the
cheapest scene to learn on.) Do the morning screen test first, ahead of
script lock and the 1a verification pass, and scope its references to the
morning's cast only — Young Printer, Old Printer, the print shop interior,
the exterior street look. Luther, Katharina and Lufft do not appear in the
morning and their stills can wait for lock. Of these, the **print shop
interior is the harder consistency problem than any face**: it carries 9 of
14 shots — one interior fewer out of thirteen since M-10 was cut, 5 Sep 2026 —
and the checklist pins two presses, counting desk, long oak table
and window placement.
- Reference stills. Generate per Services — Reference stills, iterate until approved. Working files live in `refs/` (gitignored, mirrored to the bucket by `python3 assets.py push refs`); the winner is marked with `assets.py approve` and both show on `public/refs.html`, rejects included — judging a take means seeing what it beat.
- [ ] Luther
- [ ] Katharina
- [x] Lufft *(`lufft-03`, a three-view sheet on grey — frontal, back, ¾ face close-up in one image — approved 12 Aug 2026, in the house style of the Old Printer, Young Printer and Apprentice sheets. Generated by `seedream_v5_pro` at 21:9 / 2k as a single job, conditioned on the 7 Aug sheet for identity, so it is referenceable whole by `seedance_2_5`: the 7 Aug sheet was built by `nano_banana_pro` from the 4 Aug plaster-wall takes and then had its dividers recut locally to `old-printer`'s 43 px at rgb(182,182,182), which trimmed it to 3061x1344 and cost it its job match entirely — not even panel by panel. That sheet is superseded, not deleted: `refs/char/lufft.png` stays on disk and its bucket record stands. Costume canon unchanged: flat black cap, cream linen shirt with the sleeves rolled above the elbow, pale tan canvas apron to the knee — the mid-tan register that is his and not the Old Printer's dark leather — dark blue-grey hose, brown turnshoes; ink to the elbow on both forearms in every view. Canon change, Mabel 12 Aug 2026, decided from this take: the ¾ close-up is now shot from his left — head turned toward his own right, face angling to the left edge of frame, right ear hidden — and framed tighter, close to a near-profile. It is the same side as the Apprentice's close-up and the opposite of the Old Printer's and the Young Printer's, and it supersedes the 7 Aug close-up from his right. That panel remains the face anchor for any future view. Seen and accepted alongside it: the apron ink reads as a broader, wetter stain than the old dry spatter, and the shoulder smudge of the 7 Aug close-up does not carry. Rejects `lufft-01` (close-up on the retired side, from his right) and `lufft-02` kept. No seated or in-motion still exists.)*
- [x] Old Printer *(`old-printer`, a three-view sheet on grey — frontal, back, ¾ face close-up in one image; approved 7 Aug 2026, replacing the five-view set of 3–4 Aug at Mabel's direction. The desk view `old-printer-desk` was kept alongside it at first, then unapproved and deleted 7 Aug at Mabel's direction, so the sheet is now the sole reference and no seated still exists — M-05 and M-15 are both built on the ledger pose and must carry it in the prompt. Off-axis views via nano_banana_pro, flux_2 copies the ref pose; forehead mark canon: his right forehead, left side clean)*
- [x] Young Printer *(`young-printer-smear-08`, a three-view sheet on grey — frontal, back, ¾ face close-up from his right in one image — approved 12 Aug 2026, the second approval that day and the third sheet in the lineage. Generated by `seedream_v5_pro` at 21:9 / 2k as a single job (`5b5992d2`) conditioned on the first 12 Aug sheet (`524a6cb5`), so it is referenceable whole by `seedance_2_5` — confirmed byte-exact against the job history. It supersedes both the 7 Aug sheet (locally recut, so it matched no job — the defect that forced this whole regeneration) and the first 12 Aug sheet, which stays on disk as `refs/char/young-printer.png`; retiring that record is a separate call and has not been made. The homelier face and costume carry over: pox pitting, heavy brow, brown leather jerkin front-laced and belted over a rolled-sleeve cream linen shirt. The close-up panel remains the face anchor for any future view. Known defect, seen and accepted at Mabel's direction: the two face panels disagree about the ink. The close-up carries the lampblack on his right cheekbone, which is canon; the frontal panel puts its dominant mark on his left, with only faint specks on his right. This is the same defect the superseded sheet had — heavier mark, same wrong cheek — so the approval accepts it rather than fixing it. **Canon is unchanged: his right cheekbone, left cheek clean**, as settled for the Apprentice on 7 Aug. Finding behind the acceptance — ten takes failed to fix it: `young-printer-smear-01`–`07` (`nano_banana_pro` edits) and `08`–`10` (Seedream regenerations) were all conditioned on a reference whose frontal panel carries the mark on his left, and all ten reproduced that side regardless of wording, including prompts naming the side in viewer terms and pinning it panel by panel. Three text-only Seedream takes with no image reference (`young-printer-fresh-01`–`03`, culled from disk but still served at their `result_url`) got it right 3 for 3, at the cost of a visibly different and homelier face — jug ears, heavier pox pitting — which was not adopted. So **the reference image overrides the prompt on which side a mark sits**; a directional fix has to change the reference, not the words. Mirroring the reference is not the answer either — tried on the Old Printer turnaround 4 Aug, it drifted the frontal instead. Still no in-motion still for the tracking shots M-01/02/06/09/11 — the running view was retired 4 Aug and has never been replaced.)*
- [x] Apprentice *(`apprentice`, a three-view sheet on grey — frontal, back, ¾ face close-up in one image — approved 7 Aug 2026, replacing the hero-plus-two-views set of 4 Aug; ~15, copper-red crop and freckles to contrast the Young Printer. Smudge canon, settled 7 Aug 2026 from this sheet: the lampblack smear is on his right cheekbone, the viewer's left, and his left cheek is clean — the retired hero's note called it the left cheekbone and that wording is superseded.)*
- [x] Print shop, day *(`shop-day-19`, approved 7 Aug 2026 — the pressman edited out of `shop-int-08` (itself the M-07 defect fix on 03), then both presses roughened to hard-used oak and each given its bar, and the room warmed by a local grade. The relit line 10–16 is retired: at 1:1 the model's relighting rounds had laid a halftone mesh over the plaster and smeared the wood grain, which no later round recovered. See the method caveat under the screen test. **Renamed from `shop-int-19` on 7 Aug 2026** at Mabel's direction so the morning room and the evening room do not share a name — same bytes and same hash, new filename and new shot id `SET-shop-day`; the ambiguous `SET-shop-interior` record is retired. `shop-int-NN` is now the retired day hunt. Open defect: judging the night still at 1:1 showed the halftone mesh is still here — `shop-int-09` is clean and `shop-day-19` is not, so the retexturing rounds 17–18 laid it down, not only the relighting rounds this entry blames. Any whole-surface `nano_banana_pro` pass can do it; judge at 1:1 after every round. **Sun put into the reference, 7 Aug 2026 — reversing the "flat reference" decision** at Mabel's call: 19 read as ambient fill from nowhere, and measured it was (G 11.9 below the R/B midpoint frame-wide, and the shaded floor as warm as the sunlit wall, so nothing said where the light came from). `shop-day-20` (moderate) and `shop-day-21` (lower, harder sun) are a local pass, not a model round — a warm light map falling off from each of the four windows, split-toned warm-into-lit and blue-grey into-shade, with the panes lifted so the glass reads as the source. Sunlit wall R−B +45.7 → +58.9, shaded floor +21.5 → −3.4. Local because both generation routes are known-bad here: relighting lays the mesh, and the one fresh roll lit this way (`shop-int-20`) returned "1361" on the press frame. The cost is the reason the reference was flat: the falloff is anchored to where these windows sit in this frame, so it argues for sun from the wrong side as M-03…M-15 reframe. 19 stays the approved `SET-shop-day`; 20 and 21 are added takes, not replacements. Full method in `shots-morning.md`)*
- [~] Print shop, candlelit night *(`shop-night-03`, 7 Aug 2026 — `nano_banana_pro` off `shop-day-19`, two rolls at 2 cr each: same room at night, layout and both presses intact, one candle on the long oak table, empty of people, no lamp or lantern anywhere. 01 kept sheets on the drying lines; 02 cleared them by regenerating from the day still with the empty lines described positively, rather than editing them out of 01 — and threw in a taller candle, which is what E-01 wants. 02 was not actually candlelit: its four windows measured luminance 76–94 against an 87 tabletop and threw a cold blue fill across the room, so the glass was the real key light. Mabel called it, 7 Aug. `shop-night-03` — the keeper, chosen 7 Aug 2026 — is a `nano_banana_pro` edit (2 cr) that names the windows as the change: panes to near-black, no glow on any surface, candle the only source with a steep falloff. See the light block in `shots-morning.md`. Only 03 survives. 01, 02 and 04–08 were deleted locally at Mabel's direction, 7 Aug 2026, after the whole line was reviewed in the gallery; the bucket copies remain under `gs://the-offer-assets/refs/set/`, so `python3 assets.py pull refs` restores any of them. What they were, so the line is not re-walked by accident: 04 was a local grade on 03 (a red-minus-blue warmth mask protecting the candle's pool, windows to 9–18 against a 75 tabletop, 4–8× where 02 read 0.9×); 05 and 06 masked the glass out of 02 and composited it back over 04 to put moonlight in the panes without fill in the room, at 36–42 and 51–60 respectively; 07 and 08 were a local texture-recovery pass over 06. That line is retired — 03 is a generation, not a grade, so none of 04–08's tone work carries over to it, and anything wanted from them has to be rebuilt on 03. The rules those rounds bought, which do outlive the files: kill the window's fill, not the window — what made 02 wrong was light on the walls and floor, not brightness in the panes, and the number to watch is the pane-to-tabletop ratio. And the night line's grades cost real texture: the same plaster patch measures 7.7 in `shop-day-19` but 2.7 by 06, which is what read as cartoonish. 03 has not been measured for this — it is one generation off the parent with no grades on it, so it should sit closer to 7.7, but check at 1:1 before approving. Still not approved: 03 inherits the parent's halftone mesh, and since the night still is generated from the day still, repairing the day still's texture is the fix for both. Approving now would lock the mesh into a second asset)*
- [ ] The study
- [ ] The courtyard
- [x] The town exterior, the look *(`town-ext-10`, approved 7 Aug 2026 — emptied of two dogs and three townspeople; edited from 09, itself edited from 04 to remove the lantern and downpipes)*
- [ ] `runner-path-01` — the road in (M-01)
- [ ] `runner-path-02` — the town gate (M-02)
- [ ] `runner-path-03` — the alley (M-09)
- [ ] `runner-path-04` — the printers' lane (M-11)
The exterior was missing from this list until 3 Aug 2026 and is not
optional: four morning shots are outdoors (M-01/02/09/11) and the
checklist requires them to read as one route — gate → alley → printers'
lane → shop door — in consistent early light. Approve one town look
(palette, architecture, morning sun) and generate the four locations from
it by reference rather than approving them separately. *(It was five shots
and four locations until the market was cut on 7 Aug 2026.)*
The stops are named `runner-path-NN` in route order (Mabel's direction,
7 Aug 2026), so the number says where in the run you are — and the four
numbers are the four exterior shots, M-01/02/09/11, one each. M-01 gets
`runner-path-01` even though it is feet on packed dirt with no architecture
in frame: the surface, light and palette under those feet open the film and
have to match the gate a cut later, and it is cheaper to pin them in a still
than to hope. This is a
different prefix from `town-ext-*` on purpose: `town-ext` is the hunt for
the look, `runner-path` are the places generated from it. Full convention,
with what each stop wants in frame, in `shots-morning.md`. The prefix is
registered in `build.py`'s `SCENES`, which counts approved sets per scene by
filename prefix — a new set slug that is not listed there is invisible on
the index.
**One approved hero still per subject, plus supporting angles generated
from it.* `seedance_2_0` conditions on an `image_references` array*, so
feeding the character still and the set still together beats either alone —
and a second or third angle of an approved set, generated from the hero,
strengthens the lock further. Cheap: `flux_2` is 1 credit a go.
Prompt characters descriptively ("a heavyset, clean-shaven scholar in his
fifties, black cap, 1539") — naming "Martin Luther" trips real-person
filters and drags in stock Luther iconography.
- [ ] A turnaround per character, not a single still. One portrait cannot hold a face that the film sees from every side: the Young Printer alone is shot as a wide tracking figure, a medium in profile, a face in a doorway and a top-down insert. Approve a hero three-quarter portrait, then generate front, profile, and one in-motion view from it with `--ref`, so the same face is on file at the angles the sheet actually asks for. All of them go into `seedance_2_0`'s `image_references` array together — the array is the identity lock, and feeding it one image wastes it. At ~1 credit a generation, a four-view turnaround for all five characters is ~20 credits.
- [ ] Keep sets out of character stills — that is a separate rule. Learned 3 Aug 2026: the first Young Printer pass was generated with the approved shop as an `--image-references` input and asked for "presses soft behind him". The reference did not carry — it returned a nineteenth-century cylinder press and a stone castle doorway, both anachronistic, in a portrait whose only job was the face. A character still's job is face and costume; a set still's job is the room. Shoot characters against a bare lime-plastered wall, and let the two meet at video time where character and set stills go into the references array together. This is about **what is behind the subject**, not about how many stills — the turnaround above still applies.
- [ ] Keep people and animals out of set stills — the converse rule. Mabel's direction, 7 Aug 2026: the subject of a set or background still is the place, not who is standing in it. Every figure a set still carries is a costume, a face and a pose the film never approved, and it inherits into video the way the wall lantern did in M-02 — `town-ext-09` handed `seedance_2_0` two loose dogs and three townspeople, `shop-int-08` a pressman at the bar. Both were emptied by instruction edit the same day (`town-ext-10`, `shop-int-09`), but the retrofit is not free: the pressman took the centre press's bar handle out with him, and getting it back ran to eleven takes — `nano_banana_pro` cannot be told which of two presses to hang a bar on (three of four rolls chose the wrong one), and the rounds that relit the room in the meantime cost the whole image its texture, so the approved still had to be rebuilt from `shop-int-09` and graded locally. Generate the sets below empty in the first place. Sets are cast at video time from the character turnarounds, not from whoever the set generator invented. Phrase it positively, per the clause below — negative lists get dropped. What to write: "the lane is empty, the town still asleep at first light" or "the shop stands quiet between shifts, the presses idle" — an empty place is a describable condition, whereas "no people, no dogs" is a list to ignore. Applies to the four runner-path stops (road in, gate, alley, lane) and to the study, courtyard and night shop, none of which exist yet. *(The market was the fourth until 7 Aug 2026; the difficulty of generating it empty is part of why that shot was the one cut.)*
- [ ] Anti-anachronism clause in every set prompt (learned 3 Aug 2026, the hard way — the first shop pass returned a cast-iron radiator, an electric desk lamp and two nineteenth-century presses, all plausible at a glance). Two rules that actually worked:
- State conditions positively; negative lists get ignored. "No lamps of any kind" was dropped twice. "Lit by daylight alone, falling through small leaded windows — every surface lit by that window light and by nothing else" worked first time. Likewise "contains only wood, leather, paper, stone and hand-forged iron" beat listing what to exclude.
- Name the period machinery part by part. "Wooden hand-press" is not enough; it took timber frame, central wooden screw, long bar handle, wooden platen, sliding carriage to stop getting an iron Albion press — in a film about printing.
Add the clause to each sheet's BOILERPLATE before its scene is generated;
the morning's BOILERPLATE currently has none, and "1539" alone does not
carry it.
- [ ] The contract prop is practical, not generated. The insert must legibly read `400 Gulden jährlich` with zero pseudo-text: hand-write it (or set printed props in a Fraktur/Schwabacher font such as UnifrakturMaguntia), print on laid paper, distress, photograph in set light. The photo doubles as an image-to-video conditioning reference. AI-generate only out-of-focus background pages. The route is proven, 19 Aug 2026 — M-03's `m03-plate-01` is this done end to end without a printer or a camera: type set in UnifrakturMaguntia from Luther's own German, the generated pseudo-blackletter lifted off a `seedream_v5_pro` sheet by a grayscale closing, the real type warped into the paper's plane and printed by taking the paper down toward black through its alpha, so the sheet's own sun and cockle come through the letters. The plate then drives the clip as `--start-image` and the type survived the render. Method, the text and its caveats: `notes/morning-sheet-history.md`. What stays open here is the contract, not the technique.
- [ ] Use **multi-reference image-to-video conditioned on the approved stills** — the main lever for character/set consistency, and now the criterion the model was picked on. Feed the character still and the approved set still into `image_references` together rather than settling for a single start frame.
What it is: a deliberately small trial run of the morning scene only —
approve its reference stills, then generate two or three of its shots — to
find out what the pipeline does wrong before ~70 shots are committed to it.
The morning is chosen because it clears no gates: it needs no script lock, no
Phase 1a verification and no voice engine (every line in it is VO or
off-camera walla). It runs on `seedance_2_0` with `generate_audio: false`.
What it is not: a model bake-off (Phase 2b), or the start of the volume
run — which waits on Seedance 2.5. What it teaches is prompt wording, QC and
sheet corrections, and those transfer to whichever model wins.
Two probe runs exist — don't conflate their files. Before the platform
decision there was a klingai.com free-credit probe (1 Aug 2026), driven
by `probes/video/morning-prompt-kit.txt`: takes for M-01 and M-04
plus QC frames (labelled `clipB_*` for the M-04 clip) in
`probes/video/morning/`, local-only — never pushed to the bucket. Its
findings fed the sheet rules: the type-case labels rendered pseudo-blackletter
(spelling nothing), and the runner's shoes came back as modern laced derbies
— period costume needs checking down to the feet. The **Higgsfield screen
test (3 Aug 2026)** is the M-03/M-02/M-07 run described here; its clips live
in the bucket under `refs/probe/`. So M-07 is uncovered by both runs.
Shots chosen to stress the no-text rule rather than avoid it, a print shop
being the worst case for it: M-03 (printed sheet is the subject),
M-07 (full shop in motion), M-02 (exterior, character and set at once).
- [x] Reference stills — shop interior, Young Printer turnaround (hero, profile, running), Old Printer hero, town exterior. 16.5 credits.
- [x] M-03 — set conditioning carried (the counting desk and ledger came through recognisably). Blackletter rendered beautifully and **spelled nothing**; hence the practical-text rule now in the sheet.
- [x] M-02 — the result the strategy rested on: character and set held simultaneously from a three-image reference array. Also proved reference defects inherit — the approved exterior's wall lantern propagated into the clip.
- [x] M-07 — the press mechanism in motion. *(Run 4 Aug 2026, two `seedance_2_0` takes, `refs/probe/M-07-01/02.mp4` — both fail QC, and the failures are the finding: (1) ending focused on the drying line makes generated pseudo-blackletter the subject, so the shot is re-specced to end soft on the pressman's hands; (2) the approved `shop-int-03` still itself carries a gear-toothed press screw, shirt lettering, a carved "Gottes Wort" beam and legible drying sheets — all inherited into both takes. Both notes folded into the sheet. Method note: instruction-editing the approved still beat re-rolling from the prompt — that is how the town exterior was fixed, and it names phenomena precisely: anything the prompt mentions, even as material description ("rain falls from open eaves"), gets rendered. **Cleanup done 4 Aug 2026 — `shop-int-08` approved, replacing 03. Method caveat:** FLUX.2 instruction edits re-synthesized this interior — it repainted walls and wood even when told to change one thing (takes 04–06), and describing what to keep made it worse ("cracked grey-white lime plaster" returned literal fieldstone). The edits that held were `nano_banana_pro`'s, both rounds surgical. Use `nano_banana_pro` for still cleanups; FLUX.2 edits worked on the simpler exterior but are not the general tool. **Second caveat, 7 Aug 2026 — surgical is not free, and light is not an edit.** Judge every edited still at 1:1 before approving it: `shop-int-16` was approved on content and turned out, magnified, to have a halftone mesh over its plaster and painterly smears where the wood grain had been. The rounds that did it were the two that relit the room. Colour and tone are a local grade — a per-channel affine matched to the target still's statistics, plus a gentle S-curve, costs nothing and cannot touch detail. Send the model after objects and geometry; never after light. Third caveat: the model cannot be aimed. Asked to put a bar on one of two presses, three of four rolls chose the other one, whatever description was used — position, neighbours, the paper on its own bed. When two similar objects are in frame, ask for the change on both; that is how 17 and 18 landed first time after 12–15 failed.)*
- [x] Fold the findings back into `shots-morning.md` before the full run. *(Done 3 Aug 2026: positive-framing anti-anachronism clause, period press description, practical-text rule, reference-still table, expanded continuity checklist.)*
- [x] Regenerate the town exterior without the lantern. *(4 Aug 2026: `town-ext-05..09`. Fresh re-rolls (05, 07) and describe-the-scene reference passes (06) all failed — new lanterns, shop signs, literal rain. What worked: two rounds of instruction-editing the approved still itself — 08 removed the lantern, 09 the downpipes; 09 also removed a downpipe defect PLAN hadn't recorded. **09 approved 4 Aug 2026*, replacing 04 in the manifest.)
- [ ] Clips defaulted to 720p — pass `--resolution` for anything kept.
Order: morning (wordless, cheapest to learn on) → evening → study → courtyard.
For each scene sheet:
- [ ] Generate per row: BOILERPLATE + the row's Action & camera text; 2–4 takes per shot into `takes/<scene>/<shot>/` (gitignored).
- [ ] QC every take against the sheet's continuity checklist and the two hard rules — frame-extraction with ffmpeg makes this checkable shot by shot.
- [ ] Select takes; record selects in a `selects.md` per scene so the edit is reproducible.
- [ ] **The coverage gate — runs before every audio generation, no exceptions.** The failure this prevents is silent: a word nobody decided about gets read in modern German, the take sounds fine, and it ships. So make the word list closed-world. Every distinct word in `the-offer-de.fountain` must be in exactly one of two committed lists:
- the pronunciation guide — word → respelling actually sent to the engine, or
- a waiver list — words explicitly checked and judged fine as-is.
Any word in neither fails the gate and blocks generation. That way a newly
written or edited line cannot reach the engine until someone has ruled on
every word in it, and the check reports which words are unclassified
rather than just failing. The waiver list is doing the real work here: it
is what turns "we didn't think about it" into "we decided it was fine."
This is production tooling, not a probe — it belongs committed alongside
`build.py`, and re-runs whenever the German script changes. *(If the
direct-ElevenLabs route is ever taken instead, the same gate applies with
the `.pls` lexicon standing in for the guide — the mechanism is identical,
only the artifact changes.)*
- [ ] Voices. All six parts (Luther, Lufft, Katharina, Old Printer, Young Printer, Apprentice) with the engine the 2a probe chose — from the spoken script, under the one pronunciation guide, each with its own designed voice. Check the six back to back before committing: generated ensembles go samey, and 11 minutes is long enough to hear it.
- [ ] Sound design. The sheets' Audio columns are the spec: the press's working rhythm and its silences, night quiet and quill scratch, the courtyard fire, and the hens the runner scatters in M-09. Source per Services — SFX — `mirelo_text_to_audio` for the press, quill and hen specifics, Freesound for real-world ambience. **The morning's exteriors need almost nothing else** (19 Aug 2026): the runner's shots carry the motif alone apart from those hens, so the footfalls, cart wheels and lane press-bleed the morning used to want are struck from this list — see `shots-morning.md`.
- [ ] Music. Period-flavored score, sparse — the script leans on silences. One cue is decided (4 Aug 2026, superseding "the morning opens without music"; tightened 19 Aug 2026): the runner motif — a catchy, driving beat on period instruments (pipe-and-tabor-class percussion; urgent, not menacing) that plays **whenever the runner is on screen, from the first frame until he arrives at the shop, and plays alone** — his shots (M-01/02/09/11) carry no diegetic sound but the hens he scatters in M-09, the interiors carry no music, and the M-12 door bang kills the cue dead out of pure music. Full spec in `shots-morning.md`, including the open item it creates: the approved 26.05 s take has to fit 17 s of runner screen time. `sonilo_music` per Services — Music, graded on the "1539 or trailer?" ear test, and **not mixed in until Higgsfield's commercial terms are confirmed**.
- [ ] Assembly in Kdenlive/MLT: cut selects in sheet order, lay in dialogue, sound design, music. Toolchain decision and its one real cost are at the end of this phase.
- [ ] English subtitles authored in the edit (never generated into the image): plain modern English carrying the true meaning of every German line — translate the sense, don't imitate the archaic phrasing. Text follows the English fountain; the track also translates the title cards.
- [ ] Title cards made in the edit, in German: opening **WITTENBERG, 1539* over M-01/M-02, and the closing card from the Lectures on Genesis*. The lectures survive in Latin, so the German card text is drafted and verified like the dialogue — the Matthew 10:8 clause in Luther's own Bible wording ("Umsonst habt ihr's empfangen, umsonst gebt es auch" — verify the 1545 orthography), the rest rendered into period German.
- [ ] A title sequence — the film's own title never appears. Noted 3 Aug 2026: the script opens on `TITLE CARD: WITTENBERG, 1539` and closes on the Genesis card. The Offer exists only on the fountain's title page, which is a document convention, not a shot. Four decisions, and they belong in the script before lock, not in the edit as an afterthought. **Worked through 7 Aug 2026 — name, language and typeface settled; where is narrowed to one slot but not picked.**
- What name — SETTLED: The Offer stands. Not by default: the title's job here is to make the audience wait for a thing that has not happened, and the film runs on exactly that (four minutes before you learn what the offer is, eleven before you learn what it was worth). A title that names the question is doing structural work. The alternatives weighed, all of which also withhold, in case this reopens: The Sum (from Lufft's "the sum on this paper" — points harder at the number that burns in the last shot, but reads thriller); The Contract (tracks the film's throughline prop — written, folded, carried, refused unopened, burned — but colder, and loses the sense of something proposed to a person); First Access (Lufft's own words, and the modern-negotiation register is the irony, but it half-explains the deal before the audience has met anyone, so the waiting stops). Ruled out as a class: anything naming the refusal (Freely, Gift of Grace, German Umsonst) — they answer the question the film is built to hold open.
- What language — SETTLED: English (Mabel's direction, 7 Aug 2026). Titles and credits are edit text, like the subtitles, so hard rule 2 does not bind them. Two consequences: the title becomes the **only English text in the film**, and the division that makes that coherent is that `WITTENBERG, 1539` and the Genesis card are the world's voice while the title is the film's. Those two cards stay German and subtitled.
- What typeface — follows from the language decision: not blackletter. A blackletter main title already read as pastiche to a modern audience; an English one is worse, because blackletter is the period German register itself. Set the title in something plain. The opening card and the Genesis card, being German, may still take the period face — which means the "pick once, apply to all three" rule from 3 Aug no longer holds. Pick two faces deliberately: one for the film's voice, one for the world's.
- Where — OPEN, narrowed to the cut out of M-15. The morning scene is already a cold open: ~80 s (M-01..M-15 as timed in the sheet), opening on feet with no context and ending on a hard cut into silence after "Next it will be the whole trade." That is the cold-open-then-title shape, and the silence is the slot. Candidates, in preference order: 1. Over the candle opening the evening scene — cut from the dead press to the single flame on the long table, the morning's crumpled letter beside it (already in the script, `INT. PRINT SHOP - EVENING`). Needs no new shot and no new location: the candlelit night shop is an existing Phase 3 reference-still item, not new scope. Says *here is the room where it gets decided*. 2. Over black, in silence. M-15 → black → title → the candle. Zero cost: no shot, no still, no fountain entry beyond the card. After eight seconds of a shop that has stopped making noise, black and silence hit hard. The plain option is not the weak one. 3. A rise over the rooftops off M-15 — hold the Old Printer's face in the silence, then drift up and out over the printers' lane, town beyond, morning smoke. Says here is the town, and doubles as the passage of hours into the evening. The only candidate that buys a new location: a town-from-above still generated from `town-ext-10` so it grades to the approved pair. 4. End title after the burn — the 3 Aug preference, kept live. The reveal structure argues for it: the card lands once the number has burned, and The Offer retroactively means what the printers offered, what Luther refused, and what Katharina held over the fire. Costs the film eleven minutes unnamed — normal for a short, strange for a feature.
Two placements were considered and dropped. **A hill overlooking the town,
the runner pausing on it:** Wittenberg sits on the flat Elbe lowland, so
the vantage is an invention needing a `DRAMATIC LICENSE:` flag for a shot
whose only job is to carry text (worth verifying, not verified); and the
pause deflates a runner who is carrying the news that Frankfurt and Basel
are gone, breaks the tightening route the continuity checklist enforces
(gate → market → alley → lane → door), and fights the runner motif, which
is built to snap back in on every exterior cut. **A pull-back out through
the shop roof mid-scene:** right instinct about where the tension is,
wrong mechanics — an impossible camera move is a register this otherwise
grounded, observational film never establishes, and it is not one
generation but a composite, since no model will hold the approved shop
interior and then produce a matching period town aerial in a single clip.
The rooftop rise at 3 above is that idea with the move made possible and
put at the scene break.
Sound is part of this choice. `shots-morning.md` kills the runner
motif dead on the M-12 door bang and never brings it back. A title landing
in the silence after M-15 is the one place a composer would naturally
return it, transformed — decide that with the placement, not after.
Whichever wins, it is a shot: it goes in `the-offer.fountain`, gets a row
on `shots-morning.md` (M-16), and the title card itself is still made in
the edit and never generated into frame.
- [ ] End credits — decide what an AI-generated film credits. Unsettled ground, so settle it explicitly rather than by default. Credits are edit text and are never generated into frame.
- The sources, because this project cites as it goes. Pettegree, Brand Luther, p. 290; Luther, Tischreden (WA TR 4, Nr. 4690; salary parallel Nr. 5151); the Lectures on Genesis for the closing card.
- The dramatic license, disclosed on screen. The fountain's Notes flag real invention — the sources record one conversation rather than four refusals, Katharina's scenes, Lufft as the delegation's spokesman, John Frederick standing in for both Electors. A film that rests on historical grounding says where it departed; `AGENTS.md`'s cite-as-you-go rule does not stop at the last shot.
- Synthetic-media disclosure. Picture, voices, music and most effects are generated. Platforms increasingly require this to be declared at upload; saying it in the credits is the honest version of that.
- Attribution actually owed — reconcile it here. Freesound is CC0 by decision (3 Aug 2026), so ambience owes nothing; any CC-BY file that slipped in owes a line. **Confirm whether Higgsfield's terms and the music model's require attribution at all** — that check is still open under Phase 2b, and the credits are where it lands.
- [ ] Color grade to unify takes (AI clips drift in color between generations). This is the checkbox the toolchain choice below actually costs something on — read it before starting the grade.
- [ ] Mix, export master + web encode.
Constraint, stated 7 Aug 2026: post-production tooling is open source. Not a
budget preference — a condition on the pick. DaVinci Resolve is proprietary
(Blackmagic Design; the free tier is gratis, not open source) and is therefore
out, including as a fallback for the grade. This constraint governs the edit,
grade, and mix only; the generation platform is a separate decision already
made in The platform above, and Higgsfield is a commercial service.
Kdenlive for the edit, MLT/`melt` as the engine underneath it. Both are
GPL/LGPL, so the open-source constraint is met and there is no licence to buy.
Correction — none of it is installed, and the 7 Aug text said it was (25 Aug
2026). That paragraph read *"installed and checked on this machine: kdenlive
25.12.3, melt 7.36.1, ffmpeg 8.0.1, Blender 5.2.0 LTS. Nothing to procure."*
Checked again on 25 Aug: no `kdenlive`, no `melt`, no `blender`, no `ardour`, no
`audacity` — not on `PATH`, not in apt, snap or flatpak, not in `/opt`. The only
post tool actually present is ffmpeg 7.0.2, a johnvansickle static build in
`~/.local/bin` dated 4 Aug, which is also not the 8.0.1 that was claimed. Nothing
downstream changes — the picks below stand on what the tools are, not on their
being to hand — but **Phase 6 opens with an install step it currently assumes is
done**, and the versions above are unverified until someone runs them. Re-check
at install time rather than trusting a line in this file.
The reason to prefer it here is not cost. *Kdenlive's project file is* MLT XML,
and `melt` renders that XML headless** — so the cut becomes a text artifact this
repository can generate and regenerate, exactly like everything else here. That
matters because of how this film is made: takes get regenerated constantly, and
in a conventional NLE every swap is a manual relink. If the assembly is built
from the sheets, swapping a take is an edit to a file and a re-render.
The pieces to make that real, when Phase 4 has produced enough to cut:
- `tools/scratch_track.py` already derives per-speech timings from `the-offer-de.fountain`. The same data plus the approved-take list in `assets.json` is most of an edit decision list.
- A generator writing MLT XML from those two sources would put assembly on the same footing as `build.py`: sources committed, output disposable.
- `melt project.mlt -consumer avformat:master.mov …` renders it without opening a GUI, so a render is a command and not an afternoon.
This is not a task for now — Phase 4 has to produce clips first. It is
recorded here so that when assembly starts, nobody hand-cuts what could have
been generated.
Where the choice bites: the grade. Unifying takes that drift between
generations is the one job a colorist-grade tool is built for, and Kdenlive's
grading UI is the weakest part of it — thin scopes, per-clip matching by eye.
The correction filters themselves are not the problem; `frei0r.colgate`,
`frei0r.levels`, `avfilter.lut3d` and `avfilter.colorlevels` are all present.
Plan A — measure it, don't eyeball it. Drift between generations of the same
setup is a measurement problem before it is a taste problem, and this is the
rare case where scripting beats a colorist's eye: the takes are supposed to
match each other, so the target is arithmetic, not judgement. ffmpeg carries
both halves — `signalstats`, `waveform`, `vectorscope` and `histogram` to
measure a reference frame per clip, `lut3d` to apply the correction the
measurement implies. Emit one LUT per clip, reference it from the MLT XML, and
the grade stays in the same generated-and-regenerable path as the cut. Nobody
has built this yet; do not assume it is cheap. But it is the version that
survives a take being regenerated, which hand-grading is not.
Plan B — Blender's compositor, for what measurement can't fix. Blender
5.2 LTS is installed and GPL. Its compositor is node-based with real scopes
(waveform, vectorscope, histogram), it reads and writes image sequences, and
it is driveable from Python — so it fits this repo the way a GUI-only tool
would not. Use it for shots that need an eye rather than a formula, and keep
Kdenlive for the cut.
Blender is not the recommendation for the edit itself. Its sequencer can cut
video, but for a dialogue short with subtitles and a music/SFX bed, Kdenlive is
the better instrument and MLT XML is the reason. Use Blender as a compositor
alongside it, not as a replacement for it.
Decide between A and B when there are real takes to look at, not before. If
both fall short, the answer is a better open-source path — not a proprietary
one; see the constraint at the top of this section.
Two smaller consequences, neither blocking. Kdenlive has a subtitle track
that exports SRT and can burn in, which covers the English-subtitle checkbox
above (sidecar for the web, burned where a platform demands it). Its audio
mixing is basic, so the final mix goes to Ardour (GPL, and also not installed
— see the correction above): a video timeline to cut against, EBU R128 metering
for the master, and stem export for dialogue, music and SFX. Reaper would be the
pragmatic alternative and is out under the constraint at the top of this section,
exactly as Resolve is.
Everything before the mix stays in ffmpeg (decision, 25 Aug 2026). Cue trims
at bar lines, SFX beds, loudness measurement and `morning-scene-*`-style
assemblies are scripted, not dragged: `loudnorm`/`ebur128` for levels,
`acrossfade`/`afade` for joins, `adelay`+`amix` for beds, `astats` and
cross-correlation for measuring what a file actually is. The reason is the one
that governs the cut — a scripted mix is regenerable when a take is re-rolled,
and a hand-dragged waveform is not. It is also what recovered the `morning-scene`
recipe when nobody had written it down. Audacity is fine as a pair of ears for
auditioning a cue; decisions do not live there.
- [ ] Watch-down against the two hard rules one last time on the finished cut.
- [ ] Upload/host; add a link or page on the drafts site if wanted.
- [ ] Tag the repo (`v1.0`) at the commit matching the released cut.
Numbers verified against vendor pages on 1 Aug 2026 — prices drift, so
re-check at signup. Each category is down to its pick (and, where one is
named, its fallback); the alternatives that were weighed are in git history
if a probe overturns a pick.
Higgsfield CLI is installed and authenticated (`higgsfield auth login`;
`plus` plan, 1,200 credits granted monthly). **It carries all four picture
picks** — `flux_2`, `nano_banana_pro`, `kling3_0`, `veo3_1` — so the picture
stack sits behind one auth instead of a `FAL_KEY` plus a Black Forest Labs
key plus a Google key. It also carries four TTS engines, a music model and a
text-to-audio model, which is what makes a single-platform production
possible at all. Companion skills are under `.agents/skills/`.
Measured with `higgsfield generate cost` on 3 Aug 2026:
| Images (per generation) | Cr | Video (per second) | Cr/s |
| `text2image_soul_v2` | 0.12 | `veo3_1_lite` | 1.0 |
| `flux_2` | 1 | `kling3_0` std, sound off | 1.5 |
| `nano_banana_pro` | 2 | `kling3_0` pro, sound off | 1.75 |
| `gpt_image_2` | 7 | `veo3_1` | 2.75 |
| | `seedance_2_0` | 4.5 |
Audio: `seed_audio` ≈ 0.3 cr per generation. At 59 speeches the entire
dialogue track costs well under 100 credits, so **audio is not a budget
line** — it is a quality and control problem only.
What this changes:
- Cost is no longer the deciding factor (decision, 3 Aug 2026): the subscription can be raised. Grade the 2b bake-off on look, likeness, light and policy friction — not on per-second price. The dollar figures below are kept for reference only.
- The cheap tier is Google's, not Kuaishou's. Veo 3.1 Lite (1.0 cr/s) undercuts Kling (1.5 cr/s), inverting the ordering the dollar prices implied. It is 720p, so it is the previz/blocking tier, not a bulk swap.
- Kling `pro` costs +17% over `std` — cheap enough to test whether it closes the gap to Veo on the money shots at well under Veo's rate.
- Reference stills are effectively free: FLUX.2 at 1 credit puts the whole Phase 3 turnaround, iterations included, in the tens of credits.
- Sound off saves 25%, not the 33% quoted below. Veo durations are 4/6/8 s only — 6 s shots fit, 5 s shots do not.
- Hazard — do not let the skills pick the model. `higgsfield-generate` defaults to GPT Image 2 (7 cr) and Seedance 2.0 (4.5 cr/s), the dearest option in each category. Pass the model explicitly on every call.
- Higgsfield is not a host (decision, 12 Aug 2026). Its results sit on unsigned CloudFront with immutable cache headers and byte ranges, so the drafts site could link them and skip the bucket's egress. It does not, for two reasons measured that day: 6 of the 60 files in `refs/` have no byte-exact job behind them — the graded and recomposited stills, two of them approved — so a Higgsfield-served gallery would lose them or show the wrong pixels under the right filename; and `generate list` caps at 100 per kind with no cursor (`--size 200` is rejected outright), with the image kind already returning exactly 100, so any file whose job ages out loses its URL for good. The bucket holds all 60 unconditionally. What we do keep is the URL as evidence: `provenance.py urls` recorded 81 of them while the window still had them, 28 for takes already deleted from disk.
A "pass" below ≈ 1,400 generated seconds (every shot × 3 takes).
Why this changed. The 1 Aug pick was Kling 3.0, chosen for **Subject
Binding** — locking a character from 2–4 reference stills, called then "the
strongest consistency feature on the market, and exactly our stated lever."
Checking the parameters Higgsfield actually exposes, on 3 Aug: it isn't
there. `kling3_0` takes `start_image` and `end_image` and nothing else, and
`veo3_1` takes only `start_image` — Veo's "3 ingredient images" is likewise
not reachable. Multi-reference conditioning survives in exactly one family:
| Model | Reference conditioning exposed |
| `seedance_2_0` / `_mini` | `image_references` + `video_references` + `audio_references` (mini: ≤9 images) |
| `minimax_h3` | `image_references` + `video_references`, 2K |
| `kling3_0` | `start_image` / `end_image` only |
| `veo3_1` | `start_image` only |
Character and set consistency across ~70 shots is this film's hardest
technical problem, and multi-reference conditioning is the stated lever for
it. That decides the category.
- `seedance_2_0` (4.5 cr/s) — the pick. The only family here whose reference conditioning matches the strategy. Also carries `genre`, resolution to 4k, and flexible durations. Set `generate_audio: false` — it defaults true and we dub. Bad: dearest per second in the catalog, which no longer decides anything (see the budget note); `seedance_2_0_mini` is the cheaper control.
- `minimax_h3` — the fallback, same lever, 2K, worth a bake-off slot.
- `kling3_0` — demoted to a look/motion contender, not a consistency tool. Flexible 3–15 s durations were its last structural claim, and that claim died with M-10 on 5 Sep 2026 — the 9 s shot it was the only fit for is not in the film. Occasionally sneaks English text artifacts into frame — which is why the screen test stresses the no-text rule rather than avoiding it.
- `veo3_1` — kept for the hero tier on the strength of its candlelight and golden-hour rendering (our two hardest looks), accepting that it conditions on a single frame. `veo3_1_lite` for previz. Clips 4/6/8 s only; expect rerolls on Luther-likeness shots.
Seedance 2.5 is the video model — Mabel's direction, 7 Aug 2026. All video
generation goes to `seedance_2_5`; `seedance_2_0` is the fallback and nothing
more. Released 31 July 2026, it generates a 30 s clip in one run with
multi-turn extension, and our longest speech is ~29 s, so that ceiling removes
the speech-splitting constraint entirely.
It reached the Higgsfield catalog on 7 Aug 2026 (this section's earlier
"not on Higgsfield yet, 2.0 and 2.0-mini only" is superseded), but three
measured facts stand between it and the direction, and each is a live task
rather than an objection:
- **Its reference conditioning is gated against uploads — and the way round it is a job reference. Solved 12 Aug 2026; do not re-litigate.** Every `mode omni_reference` job fed uploaded media fails `IP check not finished for input media`, on local paths and pre-uploaded media IDs alike, six retries over four minutes; `mode t2v` generates fine and 2.0 accepts the identical media IDs, so the fault is 2.5's gating, not our stills or the account. The gate never opened and does not need to. Pointing at the job that made a picture — `--image-references <job_id>`, type `image_job` — uploads nothing and so is never screened. The CLI could not express this when the escape routes below were tried; 1.1.22 (7 Aug 2026) can, and its help now reads "a UUID (upload id or job id) or a local file path". `gen.py` does the translation: a reference is always named as the still it is (`--ref refs/char/young-printer.png`, `#N` for one panel of a turnaround sheet) and is converted to a job id for models in `JOB_REF_MODELS`. It fails loudly when no job made the picture, because the fallback — an upload — is the one thing 2.5 will not take. The consequence for stills: a reference is only as good as its job. A still recomposited locally (the three-view sheets are assembled from panels) matches no single job and cannot be sent whole; `reindex.py --write` records the per-panel jobs that can. So a still meant as a video reference should be generated on the platform in its final form — every local edit after download costs it its referenceability. Coverage as of 12 Aug 2026: 14 of 19 stills referenceable whole, the Apprentice sheet by all three panels. The Young Printer entry here read "by panel 0 only" until 12 Aug 2026; that described the locally-recut 7 Aug sheet, and the approved `young-printer-smear-08` is referenceable whole.
- It caps at 720p (480p/720p), where 2.0 does native 1080p/4k — the model's own ceiling, not a Higgsfield throttle. Higgsfield's "2.5 — 4K" marketing means 720p plus a post-gen upscale. So the open `--resolution` box below cannot be satisfied on 2.5; anything kept at full resolution either goes through the upscaler or goes to 2.0. No longer true, 19 Aug 2026: `higgsfield model get seedance_2_5` now lists `resolution: 480p,720p,1080p`, and a 6s 1080p job prices at 54 credits against 39 at 720p. The cap has lifted; 2.5 no longer has to be the soft option, and nothing needs routing to 2.0 for resolution alone. We stay at 720p anyway (Mabel, 19 Aug 2026): 1080p costs 54 credits against 39 for the same 6s roll, and the credit line is the binding constraint, not the ceiling. So `gen.py`'s `seedance_2_5: {"resolution": "720p"}` default stands — it is now a deliberate choice rather than a cap, and 1080p is available with `--param resolution=1080p` when a specific take earns it. M-01-12 and M-01-13 were rolled at 1080p while the cap was being tested; every other M-01 take is 720p.
- Reviewers call it polished and smoother — better motion, character consistency and instruction following, but a glossier surface, which fights the gritty-realism direction head-on. Expect the "unretouched documentary, 35 mm grain" opener to have to work harder here than on 2.0, and judge the first 2.5 take against a 2.0 take of the same shot before committing the volume run to it.
- An earlier take is referenceable too (19 Aug 2026). 2.5 takes a `video_references` array, and a clip generated through `gen.py` carries a job id, so it goes on the wire as `video_job` and never meets the IP check that blocks uploads. `gen.py --vref <take.mp4>` does the translation. Tried on M-01: passing the best previous take forward carried its road, light and grade into the new one without re-describing them, which is the cheap way to hold continuity across re-rolls of the same shot. It carries the look, not the staging — the camera move asked for in the prompt was no likelier to land. A 1080p 6s job with one image and one video reference took over ten minutes, past `gen.py`'s wait: pick the result up with `higgsfield generate wait <job>`, `curl` it into `refs/takes/`, then `provenance.py sync`.
It also takes up to 50 reference items against 2.0's 9 images / 12 total, and
both models floor at 4 s (M-01 is specced at 3 s: generate at 4 s, trim in the
edit). A fourth cap, found the same day: **2.5 rejects prompts over 4000
characters** before generating anything, and shot prompts written for 2.0 run
past it (the M-01 whole-figure prompt was 4446) — trim before rolling.
Lifted, 19 Aug 2026: a 4704-character M-01 prompt generated normally, so
the cap is gone or has moved well above 4700. Don't trim a shot prompt down to
4000 on this account any more; the reason to keep one short is that the model
drops what it cannot hold, not that it refuses the job.
Why the gate will not clear by retrying (7 Aug 2026). The bullet above says
re-test before every take. Do it once a day, not once a take: the cause is known
and it is not ours. A plain 512×512 grey square, freshly uploaded, fails
`omni_reference` with the identical error — content, source and age are all
irrelevant, so there is nothing about our stills to fix. It matches open issue
[higgsfield-ai/cli#49](https://github.com/higgsfield-ai/cli/issues/49), filed
against `cinematic_studio_video_3_5`: the newer models gate on an `ip_detected`
field that only UI-created jobs populate, and media attached through the
public API's `medias` roles never receives that state, so the gate waits for a
verdict that will never arrive. Older models skip the gate entirely, which is
exactly why 2.0 takes the same uploads. Unresolved upstream, no maintainer
response, no workaround posted; 2.5 is a second affected model.
**Of the two escape routes tried on 7 Aug, the first one won — a week later,
and only because the CLI moved.* (1) Typed job references.* The API's media
schema accepts `nano_banana_job`, `flux_2_job`, `seedance_2_0_job` and the like
— pointing at the job that generated an image rather than at an upload of it,
which is the shape most likely to carry the screening state. **This is the
answer**, and the 7 Aug finding that "the CLI cannot express it: it converts
every UUID to `media_input`" was true of the CLI of that afternoon and is no
longer true of 1.1.22. The lesson is narrower than "the route is closed": a
capability blocked by the client, not the API, is worth re-testing after a
client release. (2) Calling the API directly. The gateway is
`https://fnf-api-gw.higgsfield.ai/fnf`, endpoints under `/developer/v2alpha/`
(`jobs`, `media`, `reference-elements`), and the CLI's stored OAuth token
authenticates fine — but every request returns `X-Fnf-Workspace-Id: Field
required` no matter how the header is sent, on every endpoint, because the
gateway strips inbound `X-Fnf-*` identity headers and injects its own for
trusted clients. Copying the CLI's user agent does not help. **The CLI is the
only working route into this API** — the lesson is the opposite of "stop using
the CLI." That leaves the web UI as the one untested path to 2.5 references,
since UI-created media is what populates the field in the first place.
**The groundwork laid while the gate was shut is what made the fix usable the
same afternoon it was found.** `tools/reindex.py` (which replaced the
`media_ids.py` sketch) matches every still to the job behind it by content
rather than by bytes — an average hash over the job's preview, so a re-save,
crop or rescale still resolves, and a recomposite honestly does not. It records
`job_content` and its distance in `refs/provenance.json`, deliberately apart
from `job`, which stays byte-exact. Where a sheet was assembled from panels it
records `panel_jobs`, which is what `--ref <sheet>#N` reaches. That store is
what `gen.py` reads to translate a named still into a job id, and the reason
the 7 Aug prediction — "those six have to be re-generated on the platform in
their final form" — turned out to cost one still rather than six.
Two standing consequences: `gen.py` records the job id as each file lands
(`provenance.job_by_result`, a lookup on the result URL it just downloaded, not
an inference), so a still generated now is referenceable in the very next take;
and the 18 MB turnaround is never re-uploaded — a job id is 36 characters.
- FLUX.2 Pro (Black Forest Labs API; ~$0.03/image) — the pick for character turnarounds and sets. Good: multi-reference conditioning, up to 8 images — approve one hero portrait per character, feed it back to hold the same face across angles and expressions; instruction-based editing (Kontext lineage); permissive about subjects; a full 5-character turnaround set costs a few dollars; sets gain too — day → candlelit night of the same print shop via references. Blackletter: proven, 3 Aug 2026 — the shop reference rendered *Gottes Wort, Heilige Schrift* and a ruled ledger head in clean Fraktur with no English anywhere, across three passes. Measured at 1 credit (`pro`, 2k), not $0.03. That weakens the case for routing in-frame text to Nano Banana Pro; try FLUX first and escalate only if a specific sheet defeats it. Bad: painterly period atmosphere has a ceiling (a taste call); it will quietly furnish a period room with modern objects unless the prompt forbids them positively — see the anti-anachronism clause in Phase 3.
- Nano Banana Pro (Google `gemini-3-pro-image`; ~$0.13–0.24/image) — the pick for any AI-generated in-frame text, and the identity fallback (up to 14 reference images). Good: the industry's best legible-text rendering, with explicit calligraphy support — the only credible AI route to Schwabacher. Bad: the most refusal-prone on real-person likenesses — always the descriptive prompt, never the name; priciest per image; Fraktur fidelity still needs a test.
No TTS anywhere is trained on Early New High German, so every engine gets
judged on the same two things: how hard is it to force the period forms,
and can delivery be directed.
The capability we gave up, stated plainly. None of the four exposes
phoneme/IPA input. Going through Higgsfield means going without
pronunciation dictionaries, so forcing is done by respelling the text
(Phase 2a). Direct ElevenLabs would restore IPA — `eleven_v3` is the only
ElevenLabs model that applies phoneme rules in German, and it does so from a
word-level `.pls` lexicon — but that means a second vendor, a second key, and
a second bill. Revisit only if the respelling route fails the probe.
Money does not decide this category: it buys minutes, not pronunciation
control. Pick on capability.
- `text2speech_v2` — the broadest, and the one to try first. Its `variant` field selects among `elevenlabs`, `minimax`, `seed_speech`, `vibe_voice`, `cozy_voice`, so ElevenLabs voices are reachable here. Note what is not exposed: no model choice and no dictionary — the quality is ElevenLabs, the IPA control is not. 57 voices via `hf voices list`. Character limits vary by variant (5,000 for elevenlabs, 15,000 for seed_speech) — irrelevant at our line lengths.
- `qwen_audio_tts` — the most directable: takes an `instruction` field (acting direction in words) plus `language: de`, `speech_rate`, `pitch_rate`, `seed`. The seed matters — it is the only engine here offering reproducible takes, which is worth a lot when re-cutting.
- `seed_audio` (~0.3 cr/generation) — cheap, and the only one accepting `audio_references` for voice cloning; `speech_rate` / `pitch_rate` / `loudness_rate` only.
- `inworld_text_to_speech` — carries just two German presets (Johanna, Josef) and exposes nothing but `prompt` and `voice`. Too thin for a six-part ensemble; include in the probe as a control, not a contender.
Decision rule: whichever engine holds the period forms with the least
respelling per line wins; directability breaks ties. Watch the ensemble
problem — six generated voices across 11 minutes go samey fast, so
differentiate hard and check them back to back before committing.
Registered voices — use these, don't clone again (16 Sep 2026). A clone
is a Higgsfield voice element: 40 credits once to make, then a line costs
the same as a preset on `seed_audio` (0.6 cr for his E-04 speech on 16 Sep, 0.2 for a three-word line — it scales with length). Cloning the same
man twice buys a second, slightly different voice for him, which is the drift
this exists to stop.
| Character | Voice id | Type | Cloned from |
| Old Printer | `e7fb0e6d-3da5-4350-a02d-1f5b3161d265` ("Old-Printer") | `element` | `refs/audio/voice-old-printer-evening-10s.wav` — E-04-01 speech + E-02-05 line, both evening |
A line in his voice: `hf generate create seed_audio --prompt "<German line>"
--voice_id e7fb0e6d-3da5-4350-a02d-1f5b3161d265 --voice_type element`.
Voice references — send them with every speaking take (17 Sep 2026,
Mabel). The video models take no voice id, only sound files, and a take
rolled with sound but without them invents new voices every time — which is
how the Young Printer ended up with three (evening E-03-01 ~200 Hz, morning
M-12-54 ~155 Hz, E-05-06 ~115 Hz). So **any take in which the Old Printer or
the Young Printer speaks goes out with his voice clip as an audio reference,
and with sound on**:
| Character | Voice reference (with the character sheets) | Higgsfield upload id | Made from |
| Old Printer | `refs/char/old-printer-voice.wav` (11 s) | `92945c7e-c2c5-4d70-9025-cba5ae69f3d4` (also `47a716c2…`, same bytes) | E-04-01 speech + E-02-05 line, both approved, both evening |
| Young Printer | `refs/char/young-printer-voice.wav` (9.9 s) | `0840af95-4ca3-4c1a-a9d7-7fd751e005de` | E-03-01 (evening) + M-12-54 (morning), both approved |
The same bytes stay at `refs/audio/voice-old-printer-evening-10s.wav` and
`voice-young-printer-approved-10s.wav` (hidden from the gallery, kept for the
thread attachments and the Old Printer clone's record).
How, and what it takes:
- Order is the naming. The prompt calls them AUDIO REFERENCE 1, 2 … in the order they are added, so add them in the order the prompt names them — E-05-07 used 1 = Old Printer, 2 = Young Printer.
- The prompt carries a VOICES block that says which reference is which man, that the clips are for who he sounds like only — never played and never quoted — and which men are silent. Each man's description and each spoken line names his reference again.
- Every spoken line is written out in the prompt. With no words on the page the model fills the gap from the clip: an edit that dropped the Young Printer's sentences (17 Sep) had to be put back before the roll.
- Sound on: `generate_audio true`. Through `gen.py`, `--aref` does it — `python3 tools/gen.py takes/E-05 --model seedance_2_5 --ref … --aref refs/char/old-printer-voice.wav --aref refs/char/young-printer-voice.wav` turns sound on and sends the files in that order. On the web, add the two files under audio in the same order.
- What it bought, E-05-07 (job `9203b32e`, 17 Sep): nothing replayed from either clip; the Old Printer's line scores 0.61 against his own clip and 0.45 against the Young Printer's, the Young Printer 0.70 against his own and 0.46 against the Old Printer's (ECAPA speaker similarity — his own morning and evening voices score 0.40 against each other). His voice landed nearer the morning (~140 Hz). It guides; it does not guarantee.
Tried and dropped for the Young Printer, 17 Sep: Voice Change on
M-12-54 (1 cr, web only — muddied the words and pulled the whole track down
~6 dB, door and room tone with it); a Seed Audio line in the `Archie` preset
dubbed over M-12-54 (0.5 cr — Mabel: "awful"); Qwen TTS refuses the Audio
page's presets outright. Preset voices, with gender and preview clips, are
listed only in the web voice picker (Audio → Voice Change → Change); the CLI
list carries names alone.
Not yet with a voice reference: Lufft (the E-02-05 sample shares the Old
Printer's pitch; no usable sample), Apprentice (under 2 s), Luther and
Katharina (no samples).
Prompt plus duration, nothing else to configure. The cues are sparse (3–5,
per Phase 5), so the ask is small; the flagship is now the morning's
runner motif (Phase 5) — catchy and driving, but still on period
instruments, so it takes the same ear test as everything else.
Grade it on the same ear test as before: *does this sound like 1539 or like
a trailer? Exposed solo lute is where it will fail. *Licensing is now a
question to settle, not a solved one** — the previous pick (ElevenLabs
Music) was chosen specifically for being trained on licensed data with film
use cleared, and that guarantee does not automatically transfer. Confirm
Higgsfield's commercial terms cover a published, possibly monetized film
before the cues go in the mix.
- `mirelo_text_to_audio` (prompt + duration) — the pick for the specifics no library will have: the wooden press cycle (lever creak, platen bite, paper slap) and quill scratch, iterated until the rhythm matches the edit.
- Freesound — CC0 only (decision, 3 Aug 2026) — kept, deliberately, for real-world ambience: town murmur, carts, hens, fire, footfalls, night quiet. Recorded beds still beat generated ones for anything the ear knows well. Filter to CC0 and take nothing else. Not a licence-risk judgement so much as a bookkeeping one: CC0 needs no attribution list to maintain, no per-file provenance to carry into the credits, and no re-audit if the film is later monetized. CC-BY would work legally but costs an end-credit line per file and a record of which file went where; CC-BY-NC is unusable outright. If a sound exists only under CC-BY, either generate it with `mirelo_text_to_audio` or do without.
- Warning — BBC Sound Effects: the famous free archive is RemArc-licensed — personal/educational use only, so it cannot go into this film's mix. Individual BBC sounds are commercially licensable via Pro Sound Effects if one proves irreplaceable.