Appearance
The content pipeline — from collection to research
Verification status — 6 October 2026
The complete path has not yet passed end-to-end acceptance. Later checks cover captures and repeat visits, classification review and human corrections, and the instructions actually sent for analysis. Their findings and limits supersede any implication below that “built” means fully verified. The stage description remains a dated inventory; deployed behavior and the remaining writers and readers still need checking.
This pipeline takes a piece from its source through analysis into the library. The following inventory was recorded on 17 August 2026 from an inspection of code, routing and gates. That inspection called the pipeline fully built; it did not establish the complete acceptance evidence now required.
The path recorded in August
- A source is scanned. One dispatcher decides which collector handles a link — five platforms each have their own collector (Instagram, TikTok, Twitter/X, and the Meta and LinkedIn ad libraries), and websites, newsrooms, and vendor portfolios each route to theirs.
- The piece is ingested with the facts that need no judgment: where it came from, what kind of file it is, which page surface. The channel is derived by a database rule — never guessed.
- The fifteen-stage spine runs: queued · download · normalize · extract frames · transcribe · identify music · classify · analyze · embed · index · extract scenes · validate · post-process · synthesize · classify scenes. Media-heavy stages run only for the formats that need them.
- The analyzer is format-aware: text, image, and video each get the prompt their route resolves to (see routing & prompts).
- Classifications are checked against controlled lists. Depending on the field, a writer accepts an exact match or an approved alternative, tries a close match, rejects an unknown value, or keeps it for review. Protecting human corrections is a requirement, but it is not yet consistently enforced; see the verification note below.
- Two gates decide what publishes. A confidence floor parks anything the analyzer is unsure about. Above it, the entity's automation setting decides: fully autonomous work publishes, supervised work waits for a person, and onboarding work always parks for review.
Safety controls described in August
- The master switch is a single database flag, and it is off. Nothing runs on a schedule until it is deliberately turned on.
- Every model call logs its cost to one table, so spend is always answerable with one query.
- The scene-indexing stage is manual-only by standing decision — it runs when a person approves it, never automatically.
Recorded limits and later correction findings
- Scale: the spine has processed a fraction of the video library; most videos have no moments yet. That is scheduled work, not a defect.
- One brand dominates: most of the library is a single brand's content, because collection has focused there while the model settles.
- Human corrections need stronger protection. A local replay on 3 October 2026 found two classification writers that can replace a human's saved label when the analyzer selects the same value. A failed replacement can also leave the previous automated label missing.
Verification note — 3 October 2026
The correction above is based on current database rules and the checked-out writer code, replayed in a disposable local database. It covers the medium a piece uses and its creative concept. It does not establish that a human edit has already been lost in production, or that the deployed code matches the checkout. The broader pipeline description above retains its August verification date; this focused check does not reverify every stage.
When the vocabulary changes
Changing a classification can mean analyzing affected work again. Joe accepted that cost while the product is being developed: the library should reflect the agreed meanings rather than preserve an old result merely because it already exists. Identify the affected work and what must change, then verify the result through its stored labels, search and review tools.
This earlier approval is not permission to start an unlimited run. The current spend and activation gates still apply. Human corrections must survive, and a failed replacement must not erase the previous result. Those protections need verification before broad reanalysis; accepting a rerun does not certify that they already work.
Where to look next
- Routing & prompts — how a piece gets its prompt.
- Continuous monitoring — what re-scans the sources.
- The state of the data — live counts, generated.