rules-ingestion-pipeline

byMaher DKF

I have a working backend system (PostgreSQL on Supabase, Python workers on Replit free tier) that validates and seals "rules" (small verified Python functions with formal pre/postconditions) through a 6-gate validation pipeline before permanent admission. This part is fully built and tested. I need you to design the missing piece: the automated ingestion pipeline that turns raw sources into validated rule candidates. CURRENT WORKING COMPONENTS (already built, do not redesign these): Database (Supabase Postgres, schema "sovereign"): - rule_candidates(candidate_id, revision, proposal, code_artifact, hoare_triple, domain_claim, domain_type, concept_maturity, digest, created_at) — immutable once inserted, trigger-enforced - rule_gate_receipts(receipt_id, candidate_id, revision, gate_id['G1'-'G6'], result['PASS'/'FAIL'/'INDETERMINATE'/'NOT_RUN'], measured_values jsonb, policy_digest, input_digest, checker_version, execution_verified) — append-only, no update/delete allowed - sovereign_rules(rule_id, seal_id, title, code_artifact, hoare_triple, active, activated_at) — final sealed output - policy_manifests(manifest_sha256, approved, approved_at) — versioned approved policy - policy_bounds(manifest_version, gate, metric, min_v, max_v) — numeric thresholds per gate - audit_events(event_id, event_type, event_payload, prev_hash, event_hash) — hash-chained append-only log - A single Postgres function sovereign.finalize_candidate(candidate_id, revision) is the ONLY path to seal a rule. It is idempotent, atomic, checks all 6 gate receipts are PASS against the current approved policy_digest, acquires an advisory lock, rejects exact code/title duplicates against currently active rules, writes the audit event, and returns the new rule_id. Python side (running on Replit, free tier, as manually-triggered scripts currently, NOT a persistent daemon): - engine_main.py: evaluate(candidate) — runs gates G1 (AST parse/compile bounds check), G4 (AST density/complexity/length metrics), orchestrates G2/G3/G5/G6 checks, then posts gate receipts to Supabase. - engine_g2.py, engine_g3.py, engine_g5.py, engine_g6.py: individual gate logic. - engine_seal.py: calls the finalize_candidate RPC after gates pass. - supa_client.py: talks to Supabase exclusively via REST using a service_role key (no direct Postgres connection from Python). THE GAP I NEED SOLVED: Right now every rule candidate is created MANUALLY: I personally write the Python function code, manually insert it into rule_candidates via SQL, manually run the gate scripts, manually call finalize_candidate. There is no automated path from a raw source to a candidate. I need an automated ingestion pipeline with this flow: 1. Input: a raw source — a PDF textbook/paper, a YouTube video URL, or a plain text excerpt from a course. 2. Extraction: pull text content from the source (PDF text extraction, or YouTube transcript/captions). 3. Chunking: split extracted text into bounded, ordered pieces with source location metadata (page number, timestamp, etc.), stored durably so work is never lost. 4. Rule proposal: send a chunk (or set of chunks) to an LLM, asking it to propose ONE small, pure, self-contained Python function (single entry point, no imports, no side effects, 120-2500 characters, loop-free preferred) that encodes a provable fact/rule from that text, along with a Hoare-style precondition/postcondition contract. 5. Insert the LLM's proposal as a new row in rule_candidates (respecting the exact schema above). 6. Trigger engine_main.py's evaluate() on that new candidate automatically. 7. If all 6 gates pass, automatically call finalize_candidate. 8. If any gate fails, leave the candidate as-is (already permanently recorded, never deleted) and move to the next chunk. HARD CONSTRAINTS: - Zero budget. Must work on Supabase free tier (pauses after 1 week inactivity, no automatic backups) and Replit free tier (containers sleep after inactivity, local disk is NOT persistent, no guaranteed always-on process). - I am a non-technical solo operator, currently working from a mobile phone only, no computer available. - No arbitrary code execution of LLM-generated code is allowed outside a strictly bounded, isolated process — the generated Python is only parsed/analyzed (AST), never executed with real inputs during validation. - LLM calls must go through a free-tier-compatible provider (e.g., OpenRouter free models) with explicit fallback handling if a model becomes unavailable or rate-limited, without silently degrading validation strength. - The pipeline must survive Replit container sleep/restart without losing in-progress work — nothing can rely on local disk or in-memory state as the source of truth; Postgres must be the durable source of truth at every step. - Must not require any paid service, VPS, or persistent server. QUESTIONS I NEED YOU TO ANSWER CONCRETELY, WITH SPECIFIC ENGINEERING DETAIL (not generic advice): 1. Given these free-tier constraints (especially Replit's sleep behavior and non-persistent disk), what concrete architecture do you propose to run this pipeline reliably? Should it be a cron-triggered script, a webhook-triggered function, a long-polling worker, or something else? Be specific about which free service triggers what. 2. What exact Python libraries would you use for PDF text extraction and YouTube transcript extraction, given the memory/time constraints of Replit free tier? 3. How would you design the "job queue" using only Postgres tables (no Redis, no external queue service) to track pending chunks and pending LLM proposals, so that a crashed/sleeping worker can resume safely without duplicating work or losing a chunk? 4. What exact prompt structure would you give the LLM to reliably get back a single valid, bounded Python function plus a formal contract, in a strict parseable format, minimizing hallucinated or malformed output? 5. How would you detect and handle an LLM proposing malformed, oversized, or unparseable code before it even gets inserted as a candidate — or do you insert everything and let the existing gate pipeline reject it? 6. What is your concrete step-by-step build order, prioritized for someone who must do this incrementally from a phone, testing one small working piece at a time? Please give an actual engineering plan, not a general description of "you would need a queue and a worker." I need concrete technology choices, concrete schema/table proposals for anything new needed, and a concrete first-build-step.

LandingIngestion Operator ConsoleIngestion Entry
Landing

Comments (0)

No comments yet. Be the first!

Landing design preview
Landing: View landing page
Ingestion Entry: Read entry explanation
Ingestion Entry: Enroll as new operator
Ingestion Entry: Verify returning identity
Ingestion Operator Console: View console
Ingestion Operator Console: Submit PDF source
Ingestion Operator Console: Submit YouTube source
Ingestion Operator Console: Submit text excerpt
Ingestion Operator Console: 1. Monitor pipeline state
Ingestion Operator Console: View sealed rule
Ingestion Operator Console: Inspect failed candidate
Ingestion Operator Console: Dismiss failed candidate
Ingestion Operator Console: 2. Retry parked work item
Ingestion Operator Console: View resumed pipeline state
Landing design preview
Landing: View landing page
Ingestion Entry: Read entry explanation
Ingestion Entry: Enroll as new operator
Ingestion Entry: Verify returning identity
Ingestion Operator Console: View console
Ingestion Operator Console: Submit PDF source
Ingestion Operator Console: Submit YouTube source
Ingestion Operator Console: Submit text excerpt
Ingestion Operator Console: 1. Monitor pipeline state
Ingestion Operator Console: View sealed rule
Ingestion Operator Console: Inspect failed candidate
Ingestion Operator Console: Dismiss failed candidate
Ingestion Operator Console: 2. Retry parked work item
Ingestion Operator Console: View resumed pipeline state