rules-ingestion-pipeline

byMaher DKF

I have a working backend system (PostgreSQL on Supabase, Python workers on Replit free tier) that validates and seals "rules" (small verified Python functions with formal pre/postconditions) through a 6-gate validation pipeline before permanent admission. This part is fully built and tested. I need you to design the missing piece: the automated ingestion pipeline that turns raw sources into validated rule candidates. CURRENT WORKING COMPONENTS (already built, do not redesign these): Database (Supabase Postgres, schema "sovereign"): - rule_candidates(candidate_id, revision, proposal, code_artifact, hoare_triple, domain_claim, domain_type, concept_maturity, digest, created_at) — immutable once inserted, trigger-enforced - rule_gate_receipts(receipt_id, candidate_id, revision, gate_id['G1'-'G6'], result['PASS'/'FAIL'/'INDETERMINATE'/'NOT_RUN'], measured_values jsonb, policy_digest, input_digest, checker_version, execution_verified) — append-only, no update/delete allowed - sovereign_rules(rule_id, seal_id, title, code_artifact, hoare_triple, active, activated_at) — final sealed output - policy_manifests(manifest_sha256, approved, approved_at) — versioned approved policy - policy_bounds(manifest_version, gate, metric, min_v, max_v) — numeric thresholds per gate - audit_events(event_id, event_type, event_payload, prev_hash, event_hash) — hash-chained append-only log - A single Postgres function sovereign.finalize_candidate(candidate_id, revision) is the ONLY path to seal a rule. It is idempotent, atomic, checks all 6 gate receipts are PASS against the current approved policy_digest, acquires an advisory lock, rejects exact code/title duplicates against currently active rules, writes the audit event, and returns the new rule_id. Python side (running on Replit, free tier, as manually-triggered scripts currently, NOT a persistent daemon): - engine_main.py: evaluate(candidate) — runs gates G1 (AST parse/compile bounds check), G4 (AST density/complexity/length metrics), orchestrates G2/G3/G5/G6 checks, then posts gate receipts to Supabase. - engine_g2.py, engine_g3.py, engine_g5.py, engine_g6.py: individual gate logic. - engine_seal.py: calls the finalize_candidate RPC after gates pass. - supa_client.py: talks to Supabase exclusively via REST using a service_role key (no direct Postgres connection from Python). THE GAP I NEED SOLVED: Right now every rule candidate is created MANUALLY: I personally write the Python function code, manually insert it into rule_candidates via SQL, manually run the gate scripts, manually call finalize_candidate. There is no automated path from a raw source to a candidate. I need an automated ingestion pipeline with this flow: 1. Input: a raw source — a PDF textbook/paper, a YouTube video URL, or a plain text excerpt from a course. 2. Extraction: pull text content from the source (PDF text extraction, or YouTube transcript/captions). 3. Chunking: split extracted text into bounded, ordered pieces with source location metadata (page number, timestamp, etc.), stored durably so work is never lost. 4. Rule proposal: send a chunk (or set of chunks) to an LLM, asking it to propose ONE small, pure, self-contained Python function (single entry point, no imports, no side effects, 120-2500 characters, loop-free preferred) that encodes a provable fact/rule from that text, along with a Hoare-style precondition/postcondition contract. 5. Insert the LLM's proposal as a new row in rule_candidates (respecting the exact schema above). 6. Trigger engine_main.py's evaluate() on that new candidate automatically. 7. If all 6 gates pass, automatically call finalize_candidate. 8. If any gate fails, leave the candidate as-is (already permanently recorded, never deleted) and move to the next chunk. HARD CONSTRAINTS: - Zero budget. Must work on Supabase free tier (pauses after 1 week inactivity, no automatic backups) and Replit free tier (containers sleep after inactivity, local disk is NOT persistent, no guaranteed always-on process). - I am a non-technical solo operator, currently working from a mobile phone only, no computer available. - No arbitrary code execution of LLM-generated code is allowed outside a strictly bounded, isolated process — the generated Python is only parsed/analyzed (AST), never executed with real inputs during validation. - LLM calls must go through a free-tier-compatible provider (e.g., OpenRouter free models) with explicit fallback handling if a model becomes unavailable or rate-limited, without silently degrading validation strength. - The pipeline must survive Replit container sleep/restart without losing in-progress work — nothing can rely on local disk or in-memory state as the source of truth; Postgres must be the durable source of truth at every step. - Must not require any paid service, VPS, or persistent server. QUESTIONS I NEED YOU TO ANSWER CONCRETELY, WITH SPECIFIC ENGINEERING DETAIL (not generic advice): 1. Given these free-tier constraints (especially Replit's sleep behavior and non-persistent disk), what concrete architecture do you propose to run this pipeline reliably? Should it be a cron-triggered script, a webhook-triggered function, a long-polling worker, or something else? Be specific about which free service triggers what. 2. What exact Python libraries would you use for PDF text extraction and YouTube transcript extraction, given the memory/time constraints of Replit free tier? 3. How would you design the "job queue" using only Postgres tables (no Redis, no external queue service) to track pending chunks and pending LLM proposals, so that a crashed/sleeping worker can resume safely without duplicating work or losing a chunk? 4. What exact prompt structure would you give the LLM to reliably get back a single valid, bounded Python function plus a formal contract, in a strict parseable format, minimizing hallucinated or malformed output? 5. How would you detect and handle an LLM proposing malformed, oversized, or unparseable code before it even gets inserted as a candidate — or do you insert everything and let the existing gate pipeline reject it? 6. What is your concrete step-by-step build order, prioritized for someone who must do this incrementally from a phone, testing one small working piece at a time? Please give an actual engineering plan, not a general description of "you would need a queue and a worker." I need concrete technology choices, concrete schema/table proposals for anything new needed, and a concrete first-build-step.

LandingIngestion Operator ConsoleIngestion Entry
Landing

Comments (0)

No comments yet. Be the first!

Project Tasks

42 planning tasks
#1

Generate system requirement document

1m 10s0.2 cr used
Done
#8

Retry Requirements

0.4 cr needed
Cancelled
#10

Generate system requirement document

1m 17s0.2 cr needed
Cancelled
#2

Generate personas & user flows

4m 42s0.2 cr used
Error
#7

Retry User Flows

0.4 cr used
Cancelled
#9

Generate personas & user flows

0.2 cr needed
Cancelled
#11

Generate personas & user flows

0m 21s0.2 cr needed
Done
#14

Create flow for Solo Operator

0m 20sCredits in parent
Done
#3

Design first page

0.2 cr needed
Cancelled
#15

Landing

3m 52sCredits in subtasks
Done
#27

Fix Landing review findings: Hero headline is clipped at mobile width in… (+2 more)

0m 41s0.4 cr used
Done
#18

Landing / Header

0m 19s0.85 cr used
Done
#19

Landing / Hero Statement

0m 20s0.85 cr used
Done
#20

Landing / Pipeline Diagram

0m 32s0.85 cr used
Done
#21

Landing / Gate Timetable

0m 35s0.85 cr used
Done
#22

Landing / Durability Ledger

0m 28s0.85 cr used
Done
#23

Landing / Entry Actions

0m 23s0.85 cr used
Done
#24

Landing / Footer

0m 17s0.85 cr used
Done
#25

App Header

0m 19sCredits in parent
Done
#26

App Footer

0m 17sCredits in parent
Done
#4

Design remaining pages

0.2 cr needed
Cancelled
#16

Ingestion Entry

2m 40sCredits in subtasks
Done
#33

Fix Ingestion Entry review finding: Trace descriptions are compressed into an unusably narrow…

0m 15s0.4 cr used
Done
#28

Ingestion Entry / Header

0m 0s0.85 cr used
Done
#29

Ingestion Entry / Explanation

0m 18s0.85 cr used
Done
#30

Ingestion Entry / Trace

0m 25s0.85 cr used
Done
#31

Ingestion Entry / Access Forms

0m 36s0.85 cr used
Done
#32

Ingestion Entry / Footer

0m 0s0.85 cr used
Done
#17

Ingestion Operator Console

8m 24sCredits in subtasks
Done
#44

Fix Ingestion Operator Console review findings: Candidate detail grid collapses at tablet width, making… (+3 more)

1m 35s0.4 cr used
Done
#34

Ingestion Operator Console / Console Header Slot

0m 0s0.85 cr used
Done
#35

Ingestion Operator Console / Console Stage Rail

0m 23s0.85 cr used
Done
#36

Ingestion Operator Console / Console Source Ledger

0m 48s0.85 cr used
Done
#37

Ingestion Operator Console / Console Chunk Inspector

1m 0s0.85 cr used
Done
#38

Ingestion Operator Console / Console Proposal Inspector

0m 54s0.85 cr used
Done
#39

Ingestion Operator Console / Console Candidate Inspector

0m 56s0.85 cr used
Done
#40

Ingestion Operator Console / Console Sealed Rule Inspector

0m 33s0.85 cr used
Done
#41

Ingestion Operator Console / Console Provider State Panel

1m 7s0.85 cr used
Done
#42

Ingestion Operator Console / Console Pipeline Control

0m 25s0.85 cr used
Done
#43

Ingestion Operator Console / Console Footer Slot

0m 0s0.85 cr used
Done
#5

Architecture

0.2 cr needed
Cancelled
#6

Workspace task plan

0.2 cr needed
Cancelled
Landing design preview
Landing: View landing page
Ingestion Entry: Read entry explanation
Ingestion Entry: Enroll as new operator
Ingestion Entry: Verify returning identity
Ingestion Operator Console: View console
Ingestion Operator Console: Submit PDF source
Ingestion Operator Console: Submit YouTube source
Ingestion Operator Console: Submit text excerpt
Ingestion Operator Console: 1. Monitor pipeline state
Ingestion Operator Console: View sealed rule
Ingestion Operator Console: Inspect failed candidate
Ingestion Operator Console: Dismiss failed candidate
Ingestion Operator Console: 2. Retry parked work item
Ingestion Operator Console: View resumed pipeline state
Landing design preview
Landing: View landing page
Ingestion Entry: Read entry explanation
Ingestion Entry: Enroll as new operator
Ingestion Entry: Verify returning identity
Ingestion Operator Console: View console
Ingestion Operator Console: Submit PDF source
Ingestion Operator Console: Submit YouTube source
Ingestion Operator Console: Submit text excerpt
Ingestion Operator Console: 1. Monitor pipeline state
Ingestion Operator Console: View sealed rule
Ingestion Operator Console: Inspect failed candidate
Ingestion Operator Console: Dismiss failed candidate
Ingestion Operator Console: 2. Retry parked work item
Ingestion Operator Console: View resumed pipeline state