ai-career-pulse

byKinjal bascom

ROLE: Senior full-stack + data engineer. Build "AI Career Pulse" in phases. After each phase: run tests, summarize, list open issues, STOP for my review. Never invent endpoints, params, fields or data. Log unknowns/assumptions in NOTES.md and ask me. GOALS (priority order) 1. Rank AI/ML jobs by fit to MY profile (transparent score). 2. Show my skill gaps vs demand + learning plan. 3. JD-specific interview prep. 4. Track applications. 5. Market trends dashboard (stock-style), labeled with source limits. Roles: AI Engineer, AI/ML Engineer, ML Engineer, LLM/GenAI Engineer, Data Scientist (ML), MLOps, Computer Vision, NLP Engineer. MODE: personal (default): single user, APScheduler, PostgreSQL, no Celery/Redis/SSE. Keep interfaces clean so a later product mode (auth, Celery, SSE) is additive. KEYS (backend .env only; never in frontend/logs/fixtures/git; mask in logs; ship .env.example): INDIANAPI_JOBS_KEY (job data), TYPESAFE_API_KEY (Jev decisions), LLM_API_KEY + LLM_PROVIDER=openai|anthropic (text only), optional YOUTUBE_API_KEY, GITHUB_TOKEN, alert credentials. SOURCES A) IndianAPI (job data): GET https://jobs.indianapi.in/jobs, header X-Api-Key. Confirmed param: limit, sent as STRING ("50"); int gives 422. Unconfirmed: title, location, company, experience, job_type, pagination. Response array: id, title, company, about_company, job_description, job_title, job_type, location, experience, role_and_responsibility, education_and_skills, apply_link, posted_date (ISO 8601). No salary field: show "Not available from source", never estimate. Rate limits unknown: limiter (default 1 req/2s), INDIANAPI_DAILY_CAP, retry 429/5xx max 3 with backoff+jitter, honor Retry-After. B) TypeSafe Jev (decisions only, generates NO text): POST https://api.typesafe.ai/v1/systemone, Bearer auth; use typesafe-sdk (AsyncTypeSafeClient, Choice/Score/Noul, RetryPolicy). Body {state, model, questions}. choice->choice+probabilities+confidence; score->score+probabilities+confidence; noul->probability 0-1. Pin JEV_MODEL (exact id via SDK list-models; not jev-latest). Store model version per record. First install skill: claude plugin marketplace add typesafe-ai/skills && claude plugin install typesafe@typesafe-ai. Read docs.typesafe.ai primitives, confidence, patterns (fan-out, confidence routing, composite scoring), cookbooks (rerank, entity alignment, pre-parsed extraction), jev-1.13 jaggedness; summarize limits in NOTES.md. C) LLM (text): provider-agnostic, temp 0.2, JSON-schema-validated, cached. D) Prep: YouTube Data API, Dev.to, HN Algolia, arXiv, GitHub search, RSS. Store title, URL, date, metadata, short summary in our own words; never full copyrighted text. HARD RULES: No scraping any website (no LinkedIn/Indeed/Naukri/Glassdoor). Missing field=NULL; no fake data outside fixtures. Every job shows source + apply_link. Low-confidence decisions -> "needs review", never silently accepted. Market pages show banner "Source: IndianAPI only; trends reflect this source, not the full market." STACK: Python 3.11, FastAPI, SQLAlchemy 2, Alembic, httpx, Pydantic v2, APScheduler, PostgreSQL 16 + pgvector, rapidfuzz, typesafe-sdk. Next.js 14, TypeScript, Tailwind, TanStack Query, lightweight-charts, Recharts. Docker Compose, pytest, ruff, mypy, GitHub Actions. CONFIG (config/*.yaml): profile (skills {name, level 1-5}, years, domains, cities, remote pref, target roles, min seniority, deal-breakers, resume_text), roles (canonical+description+synonyms+regex), skills (id, name, description), locations (aliases e.g. Bengaluru->Bangalore), thresholds, fit_weights, schedule, sources. PHASE 0 PROBES (stop and report) 1. probe_indianapi.py: limit="5", then test each unconfirmed param and pagination; record status, count, whether results truly filtered; check posted_date freshness, id stability, max limit. Save to tests/fixtures/indianapi/. 2. probe_jev.py: 10 fixture JDs through the enrichment questions; measure latency, tokens, confidence spread, batched vs split calls, state-length limit. Save to tests/fixtures/jev/. 3. NOTES.md: results, query plan, batch size. PIPELINE: fetch -> normalize -> loose regex pre-filter -> dedup -> Jev enrich -> gate -> upsert -> embed -> fit -> metrics -> alerts. - Query plan: if filters work, loop roles x cities within cap; else fetch max and filter locally. - Normalize: source_job_id=id; title=job_title or title; keep JD sections separate; location->city/state/country; posted_at=parse(posted_date); raw JSON in jobs.raw. - Dedup: sha256(norm company+title+city). rapidfuzz title >=85, same company, 30 days -> Jev score same-posting 0/1/2: 2 merge (keep earliest posted_at, all links), 1 review. - Unseen in 3 full syncs -> inactive. JEV ENRICHMENT (batched per job; state = title, company, location, experience, job_type, JD sections, truncated per Phase 0) - is_ai_role (noul): primary work is AI/ML/DS/LLM/CV/NLP/MLOps engineering. - role (choice): roles.yaml + other. - seniority (score): fresher, junior 0-2y, mid 2-5y, senior 5-8y, lead 8y+. - work_mode (choice): onsite|hybrid|remote|unspecified. - skill_<id> (noul each, chunked): job requires/prefers <skill>. - experience (choice): regex-extracted spans + none; code converts to min/max. - domain (choice): fintech, real estate, hospitality, healthcare, e-commerce, HR tech, document AI, manufacturing, SaaS, other. Gating: is_ai_role >=0.7 keep, 0.4-0.7 review, <0.4 drop (logged). Skill >=0.6 attach with prob. Low confidence -> NULL + needs_review (or LLM fallback if enabled). Cache by sha256(state+question_set_version+model). API failure -> status failed, retry next run, never block ingestion. CAREER MODULE Fit (one Jev call/job, state = JD + profile summary): - skill_match score 0-3, experience_match score 0-2 (below/match/above), domain_match noul, growth_potential score 0-2 (scope beyond my level), red_flags noul (unrealistic/vague/mislabeled). - Code: location/remote prefs, deal-breakers (hard filter), min seniority. - fit 0-100 = weighted sum (fit_weights). Store components+confidence; UI shows "why this score". Recompute on profile change. Skill gap: per skill 30/90-day trend, senior-vs-mid share, share in my top-fit jobs, my level. Gap = high on these AND my level <=2. /growth: top 5 gaps, ranked resources, 4-week plan (LLM, grounded only in matched resources, cite URLs). Prep per job: shortlist top 30 by 0.5*cosine+0.5*skill overlap; Jev rerank score per (JD, resource) pair (not/somewhat/very useful); keep 10. POST /jobs/{id}/interview-questions: LLM, 15 questions (technical, ML system design, behavioral), tag my gap skills, cached. Tracker: saved, applied, interview, offer, rejected, withdrawn; dates, notes, next action; funnel chart. Alerts: daily digest (email|telegram) of new jobs fit >= ALERT_MIN_FIT + top 3 rising gap skills. MARKET METRICS: per skill/role/city: active, new today, delta% vs yesterday and 7-day avg. AI Job Index = active AI jobs, base 100 day 1. Volume swing >50% -> "source anomaly", excluded from deltas. Jev model change -> chart marker. "Insufficient history" until 7 days. SCHEDULE (Asia/Kolkata): poll every 2h 08:00-22:00; full sync 02:00; prep 03:00; metrics+fit+gap 04:00 and after runs; digest 08:30. Log ingestion_runs and api_usage (service, calls, tokens). DATA MODEL: companies, jobs (role, seniority, work_mode, domain, ai_role_prob, exp_min/max, confs, needs_review, model_version, raw, embedding, is_active), job_sources, skills, job_skills(prob), job_fit(components JSONB), skill_gap_daily, prep_resources, resource_skills, job_prep_rank, interview_question_sets, applications, review_queue, daily_metrics, ingestion_runs, api_usage, decision_cache, alerts_sent. Unique (source, source_job_id); indexes on posted_at, role, city, is_active, dedup_hash, fit; pgvector. API /api/v1: GET /jobs (filters role, skill, city, work_mode, seniority, min_fit, posted_within_days, q; sort fit|date), GET /jobs/{id}, GET /jobs/{id}/prep, POST /jobs/{id}/interview-questions, GET/PUT /profile, GET /growth/gaps, GET /growth/plan/{skill}, GET/POST/PATCH /applications, GET /market/{index,ticker,skills,locations}, GET /prep, GET /admin/{health,runs,api-usage,review-queue}, POST /admin/review/{id} (saved as labels). FRONTEND: / my top-fit new jobs + gaps; /jobs; /jobs/[id] (JD, fit breakdown, matched/missing skills, apply, prep, questions, track); /growth; /applications; /market (ticker tape, index chart, gainers/decliners, banner); /skills/[name]; /prep; /profile; /admin. Responsive, dark mode, loading/empty/error states. TESTS: fixtures only, no live calls in CI. Cover normalize, string limit, error codes, Jev parsing/gating/cache/fallback, dedup, experience regex, fit, gap, metrics, anomaly guard, alerts. Eval script: I label 50 JDs (is_ai_role, role, skills, fit 1-5); report precision/recall per threshold and fit correlation; tune thresholds and weights. Coverage >=80%. PHASES: P0 probes; P1 skeleton, DB, configs; P2 IndianAPI connector, dedup, scheduler; P3 Jev enrichment, gating, review queue; P4 fit + alerts; P5 prep + rerank + questions; P6 gap + plans + tracker; P7 market metrics; P8 frontend; P9 eval, CI, README. OUTPUT RULES: list files before each phase; missing key -> skip service with log; ask before adding any new service or dependency.

/Login/growthSign Up/jobs/[id]
/

Comments (0)

No comments yet. Be the first!

System Requirements

Page 1 of 23

System Requirements Document for ai-career-pulse

1. Introduction

AI Career Pulse is a personal, single-user AI/ML job intelligence instrument. It ingests AI/ML job postings from the IndianAPI jobs endpoint, enriches each posting with structured decisions from TypeSafe Jev, ranks every job against the owner's own profile with a transparent 0–100 fit score, exposes skill gaps versus observed market demand with a grounded learning plan, generates JD-specific interview preparation, tracks applications through a funnel, and renders a stock-style market trends dashboard that is explicitly labeled with its source limits.

The product is built in phases. After each phase the builder runs tests, summarizes results, lists open issues, and stops for the owner's review. Unknowns and assumptions are logged in NOTES.md and raised as questions; endpoints, parameters, fields, and data are never invented.

The audience is a single technical power user — an AI/ML candidate who reads tables, scores, and charts all day — operating the tool for their own job search. The interface is a warm graphite instrument panel, not a marketing surface.

Page 2 of 23

2. System Overview

AI Career Pulse runs in personal mode by default: one user, APScheduler for scheduling, PostgreSQL 16 with pgvector for storage and embeddings. There is no Celery, no Redis, and no SSE. Interfaces are kept clean so that a later product mode (auth, Celery, SSE) is additive rather than a rewrite.

Current actors:

  • Job Seeker (AI/ML candidate) — the single personal-mode user who maintains the profile, reviews fit breakdowns, resolves needs-review items, consumes the daily digest, prepares for interviews, and tracks applications.
  • Operator / Reviewer (admin) — the same owner acting operationally: monitoring ingestion runs, API usage, health, and the review queue; labeling review-queue items; running Phase 0 probes; reading NOTES.md; and reviewing each phase's test results, summary, and open issues.

External and system actors (not personas): the IndianAPI jobs service (job data), TypeSafe Jev (decisions only, generates no text), the configured LLM provider (text only), prep-resource providers (YouTube Data API, Dev.to, HN Algolia, arXiv, GitHub search, RSS), the email or Telegram delivery channel for the daily digest, and the internal APScheduler-driven pipeline.

Accepted behavior spans five goals in priority order: (1) rank AI/ML jobs by fit to the owner's profile with a transparent score; (2) show skill gaps versus demand plus a learning plan; (3) JD-specific interview prep; (4) track applications; (5) a stock-style market trends dashboard labeled with source limits.

Narrow exclusions and boundaries:

  • No scraping of any website, including LinkedIn, Indeed, Naukri, and Glassdoor.
  • No salary field exists in the IndianAPI response; the UI shows "Not available from source" and never estimates salary.
  • Jev is decisions-only and generates no text; the LLM is text-only.
  • Missing fields are stored as NULL; no fake data exists outside test fixtures.
  • Market pages always carry the banner: "Source: IndianAPI only; trends reflect this source, not the full market."
  • Low-confidence decisions become "needs review" and are never silently accepted.
Page 3 of 23

2a. Product Interpretation and Delivery Boundary

Delivery ownership. AI Career Pulse is a first-party application with its own backend (FastAPI, /api/v1) and its own Next.js frontend. All job ranking, fit scoring, gap analysis, prep ranking, application tracking, and market metrics are computed and rendered by the application itself. Job data is fetched from the IndianAPI jobs endpoint under the owner's key; structured decisions come from TypeSafe Jev; free text (learning plans, interview questions) comes from the configured LLM provider. Prep resources are discovered through the listed providers and stored as metadata plus a short summary written in our own words — never as full copyrighted text.

Access ownership. The application owns identity for its own durable state. Because the profile, fit history, applications, review labels, and operator functions must remain bound to the correct participant and be resumable, the application provides self-service enrollment and returning verification. The public entry surface (/) is anonymously reachable and explains the product before any protected work; the protected destinations (/profile, /jobs, /jobs/[id], /growth, /applications, /market, /skills/[name], /prep, /admin) require an established identity. Operator functions (review queue, health, runs, API usage, phase probe reporting) require operator authorization. Application identity and session continuity do not by themselves establish differentiated permissions beyond the operator boundary described here.

Current vs future boundary. Current scope is personal mode: single user, APScheduler, PostgreSQL, no Celery/Redis/SSE. A later product mode (auth, Celery, SSE) is explicitly a future horizon and is not part of current pages or acceptance. The daily digest is delivered through an external email or Telegram channel; delivery configuration is required for that capability, and delivery itself remains provider/external-channel owned.

Source limits. Market trends reflect only the IndianAPI source. Volume swings greater than 50% are flagged as "source anomaly" and excluded from deltas. Charts show "Insufficient history" until seven days of data exist. A Jev model change is marked on the chart.

Page 4 of 23

2b. Source Content Inventory

Prep resources are collected from the following content sources. For each resource the application stores title, URL, date, metadata, and a short summary written in our own words; full copyrighted text is never stored.

  • YouTube Data API — video resources; optional YOUTUBE_API_KEY.
  • Dev.to — article resources.
  • HN Algolia — discussion and link resources.
  • arXiv — paper resources.
  • GitHub search — repository resources; optional GITHUB_TOKEN.
  • RSS — feed resources.

Job content is sourced from the IndianAPI jobs endpoint with the documented response fields: id, title, company, about_company, job_description, job_title, job_type, location, experience, role_and_responsibility, education_and_skills, apply_link, posted_date (ISO 8601). There is no salary field.

2c. Page Content and Component Coverage

/

  • Information/state: Anonymous first impression of AI Career Pulse. Full-bleed graphite hero band with hairline top and bottom rules. Left: flush-left oversized headline "YOUR NEXT ROLE, SCORED." with "SCORED" in tangerine. Directly beneath, a single ruled line of live data in monospace: AI Job Index value, active AI job count, new-today count, and delta versus 7-day average. Right: the top three fit-ranked jobs as three hairline-ruled rows — company, role, city on the left; an oversized tabular fit numeral right-aligned in acid lime with a 3px probability bar beneath it. Below the hero: the owner's top-fit new jobs and current top gaps.
  • Primary actions: Enter the product (proceed to Login or Sign Up); open a listed job.
  • Supporting actions: Read the source-limits strip; scan the ticker tape.
  • Domain entities: jobs, job_fit, daily_metrics, skill_gap_daily.
  • Component responsibilities: hero band (headline + live data line + top-three fit rows); full-bleed ticker tape; top-fit new jobs list; top gaps list; source-limits strip.
  • States: loading (skeleton rules); empty (no jobs ingested yet — show the source-limits strip and a neutral "no data yet" line, never fabricated numbers); success (hero data line and top-three rows populated); error (source unavailable — show the strip and a muted error line, keep the entry actions available); recovery (retry on next data refresh).
Page 5 of 23

Login

  • Information/state: Returning verification for the owner. Minimal graphite panel with a single hairline-ruled form.
  • Primary actions: Submit credentials to establish a session.
  • Supporting actions: Navigate to Sign Up; return to /.
  • Domain entities: identity/session.
  • Component responsibilities: credential form; error region; link to Sign Up.
  • States: loading (submit disabled, inline progress); empty (n/a); success (redirect to the intended protected destination); error (invalid credentials — inline message, no field echo); recovery (retry; link to Sign Up).

Sign Up

  • Information/state: Self-service enrollment for the Job Seeker. No invitation or provisioning boundary exists, so enrollment is self-service.
  • Primary actions: Create the owner's identity.
  • Supporting actions: Navigate to Login; return to /.
  • Domain entities: identity.
  • Component responsibilities: enrollment form; validation region; link to Login.
  • States: loading (submit disabled); empty (n/a); success (identity created, proceed to profile setup); error (validation or conflict — inline message); recovery (correct and resubmit).

/profile

  • Information/state: The owner's profile used in fit computation: skills with level 1–5, years, domains, cities, remote preference, target roles, minimum seniority, deal-breakers, and resume text. Includes a schematic diagram showing how the score is composed as weighted bars labeled skill_match, experience_match, domain_match, growth_potential, red_flags.
  • Primary actions: Edit and save profile fields; trigger recomputation of fit on profile change.
  • Supporting actions: Review the score-composition schematic; inspect current fit weights.
  • Domain entities: profile, fit_weights, skills, locations.
  • Component responsibilities: profile form; skill-level editor; score-composition schematic; save/recompute control.
  • States: loading (form skeleton); empty (first use — prompt to complete profile); success (saved, recompute queued); error (validation or save failure — inline message, unsaved changes preserved); recovery (retry save).
Page 6 of 23

/jobs

  • Information/state: Ranked, filterable job list. Filters: role, skill, city, work_mode, seniority, min_fit, posted_within_days, and free-text q; sort by fit or date. Each row shows company, role, city, posted date, source, apply link, and the fit numeral right-aligned in a fixed-width column with a 3px probability bar beneath it.
  • Primary actions: Apply filters; change sort; open a job.
  • Supporting actions: Clear filters; scan the ticker tape.
  • Domain entities: jobs, job_fit, job_skills, companies, job_sources.
  • Component responsibilities: filter bar; sort control; ruled job table; fit numeral column; source-limits strip where market context appears.
  • States: loading (row skeletons); empty (no jobs match — show active filters and a clear-filters action); success (rows populated); error (fetch failure — inline error with retry); recovery (retry; adjust filters).

/jobs/[id]

  • Information/state: Two-column split — JD and fit breakdown on the left at 62%, matched/missing skills, prep resources, and tracker controls in a sticky right column at 38%. Shows the JD sections separately, the fit breakdown with components and confidence, the "why this score" explanation, matched and missing skills, the source and apply link, and the salary line reading "Not available from source".
  • Primary actions: Open the apply link; generate interview questions; save/track the application; open prep resources.
  • Supporting actions: Review fit components; review matched/missing skills; read the source-limits strip.
  • Domain entities: jobs, job_fit, job_skills, prep_resources, job_prep_rank, interview_question_sets, applications.
  • Component responsibilities: JD panel; fit-breakdown panel with component bars and confidence; matched/missing skills list; sticky right rail with prep resources and tracker controls; apply action; interview-questions panel.
  • States: loading (split skeleton); empty (no prep resources yet — show a neutral line and a generate action); success (all panels populated); error (job fetch or question generation failure — inline error, apply link still available); recovery (retry generation; retry fetch).

/growth

  • Information/state: One continuous ruled list of skill-gap rows. Each row: skill name, a horizontal bar split into a tangerine filled portion (my level 1–5) and an acid-lime demand portion (30-day market share), with the numeric gap right-aligned in tabular figures. Shows the top 5 gaps, ranked resources, and a 4-week plan grounded only in matched resources with cited URLs.
  • Primary actions: Open a skill detail; open a ranked resource; read the 4-week plan.
  • Supporting actions: Scan the ticker tape; read the source-limits strip.
  • Domain entities: skill_gap_daily, skills, prep_resources, resource_skills.
  • Component responsibilities: ruled gap list; split demand/level bars; top-5 gaps block; ranked resources list; 4-week plan panel with citations.
  • States: loading (row skeletons); empty (insufficient history — show "Insufficient history" until 7 days); success (rows and plan populated); error (plan generation failure — show resources without plan and an inline error); recovery (retry plan generation).
Page 7 of 23

/applications

  • Information/state: Application tracker with statuses saved, applied, interview, offer, rejected, withdrawn; dates, notes, and next action per application; funnel chart.
  • Primary actions: Create an application; change status; edit dates, notes, and next action.
  • Supporting actions: Open the linked job; read the funnel chart.
  • Domain entities: applications, jobs.
  • Component responsibilities: application table; status control; notes and next-action editor; funnel chart.
  • States: loading (table skeleton); empty (no applications — prompt to track from a job); success (rows and funnel populated); error (save failure — inline error, edits preserved); recovery (retry save).

/market

  • Information/state: Stock-style market dashboard: ticker tape, AI Job Index chart (base 100 on day 1), gainers/decliners, and the permanent source-limits banner. Volume swings greater than 50% are marked "source anomaly" and excluded from deltas; Jev model changes are marked on the chart; "Insufficient history" appears until 7 days.
  • Primary actions: Read the index chart; scan gainers/decliners; open a skill or location context.
  • Supporting actions: Read the source-limits banner; hover the ticker tape to pause it.
  • Domain entities: daily_metrics, skill_gap_daily, jobs.
  • Component responsibilities: full-bleed ticker tape; index chart with anomaly and model-change markers; gainers/decliners table; permanent source-limits strip.
  • States: loading (chart skeleton); empty ("Insufficient history" until 7 days); success (chart and tables populated); error (metrics fetch failure — inline error, banner remains visible); recovery (retry on next metrics run).

/skills/[name]

  • Information/state: Named-skill context within the growth experience: 30/90-day trend, senior-versus-mid share, share in the owner's top-fit jobs, the owner's level, and the computed gap.
  • Primary actions: Read the trend and share breakdown; open ranked resources for the skill.
  • Supporting actions: Navigate back to /growth; read the source-limits strip.
  • Domain entities: skills, skill_gap_daily, job_skills, prep_resources, resource_skills.
  • Component responsibilities: trend chart; senior/mid share bars; top-fit share indicator; level indicator; ranked resource list.
  • States: loading (chart skeleton); empty (insufficient history — show "Insufficient history"); success (all panels populated); error (fetch failure — inline error with retry); recovery (retry).
Page 8 of 23

/prep

  • Information/state: Preparation resources and job-specific ranked prep. Shows the shortlist of top 30 candidates by 0.5*cosine + 0.5*skill overlap and the Jev-reranked top 10 per job, with rerank scores (not/somewhat/very useful).
  • Primary actions: Open a resource; open a job's prep; regenerate questions from a job.
  • Supporting actions: Filter by skill; read the source-limits strip.
  • Domain entities: prep_resources, resource_skills, job_prep_rank, interview_question_sets.
  • Component responsibilities: ranked prep list; rerank score column; resource metadata rows; link to job detail.
  • States: loading (row skeletons); empty (no prep yet — prompt to run prep); success (ranked list populated); error (fetch failure — inline error with retry); recovery (retry).

/admin

  • Information/state: Operational monitoring and review-queue labeling: health, ingestion runs, API usage (service, calls, tokens), and the review queue. Also surfaces Phase 0 probe reporting. Monospace numerals for index values, probabilities, and API latency.
  • Primary actions: Label a review-queue item (saved as labels); inspect runs and API usage.
  • Supporting actions: Read health status; read probe results.
  • Domain entities: ingestion_runs, api_usage, review_queue, decision_cache.
  • Component responsibilities: health panel; runs table; API-usage table; review-queue list with labeling controls; probe-report panel.
  • States: loading (panel skeletons); empty (no review items — neutral line); success (panels populated); error (fetch failure — inline error with retry); recovery (retry).
Page 9 of 23

3. Functional Requirements

Each requirement is a distinct story point with provenance, lifecycle facts, and observable acceptance.

FR-1 — Phased build with review stops (explicit) As the Operator / Reviewer, I should have the build proceed in phases (P0 probes; P1 skeleton, DB, configs; P2 IndianAPI connector, dedup, scheduler; P3 Jev enrichment, gating, review queue; P4 fit + alerts; P5 prep + rerank + questions; P6 gap + plans + tracker; P7 market metrics; P8 frontend; P9 eval, CI, README), so that after each phase tests run, results are summarized, open issues are listed, and the build stops for my review.

  • Trigger: phase completion. Observable result: test run, summary, open-issues list, and a stop. Failure/recovery: failing tests are reported as open issues before the stop. Continuation: next phase begins only after review.

FR-2 — No invented endpoints, params, fields, or data (explicit) As the Operator / Reviewer, I should have unknowns and assumptions logged in NOTES.md and raised as questions, so that no endpoint, parameter, field, or data value is invented.

  • Trigger: any unknown encountered. Observable result: entry in NOTES.md and a question. Failure/recovery: unresolved unknowns block the affected work rather than being guessed. Continuation: work resumes after an answer.

FR-3 — Rank AI/ML jobs by fit (explicit) As the Job Seeker, I should see AI/ML jobs ranked by fit to my profile with a transparent 0–100 score, so that I can prioritize the best-matched roles.

  • Trigger: pipeline run or profile change. Observable result: ranked list with fit numerals and component breakdown. Failure/recovery: missing components leave the score incomplete and flagged. Continuation: recompute on profile change.

FR-4 — Skill gaps versus demand plus learning plan (explicit) As the Job Seeker, I should see my skill gaps versus observed demand plus a learning plan, so that I can close the highest-value gaps.

  • Trigger: metrics/gap run. Observable result: top 5 gaps with 30/90-day trend, senior-vs-mid share, share in my top-fit jobs, and my level; a 4-week plan grounded only in matched resources with cited URLs. Failure/recovery: insufficient history shows "Insufficient history" until 7 days. Continuation: plan regenerates as data accumulates.

FR-5 — JD-specific interview prep (explicit) As the Job Seeker, I should get JD-specific interview preparation, so that I can prepare for a specific role.

  • Trigger: opening a job and requesting questions. Observable result: 15 questions (technical, ML system design, behavioral) tagged with my gap skills, cached. Failure/recovery: generation failure shows an inline error and preserves the apply link. Continuation: retry generation.

FR-6 — Track applications (explicit) As the Job Seeker, I should track applications through saved, applied, interview, offer, rejected, and withdrawn with dates, notes, and next action, so that I can manage my pipeline.

  • Trigger: tracking action from a job. Observable result: application row and funnel chart update. Failure/recovery: save failure preserves edits. Continuation: status changes over time.

FR-7 — Market trends dashboard (explicit) As the Job Seeker, I should see a stock-style market trends dashboard labeled with source limits, so that I can read market direction without over-trusting the source.

  • Trigger: metrics run. Observable result: ticker tape, AI Job Index chart, gainers/decliners, and the permanent banner. Failure/recovery: anomalies excluded from deltas; "Insufficient history" until 7 days. Continuation: daily updates.

FR-8 — Target roles (explicit) As the Job Seeker, I should have jobs classified into AI Engineer, AI/ML Engineer, ML Engineer, LLM/GenAI Engineer, Data Scientist (ML), MLOps, Computer Vision, and NLP Engineer, so that ranking matches my target roles.

  • Trigger: Jev enrichment. Observable result: role assigned from roles.yaml plus other. Failure/recovery: low confidence yields NULL and needs_review. Continuation: re-enrichment on model or question-set change.

FR-9 — Personal mode (explicit) As the Operator / Reviewer, I should run in personal mode (single user, APScheduler, PostgreSQL, no Celery/Redis/SSE) with clean interfaces, so that a later product mode (auth, Celery, SSE) is additive.

  • Trigger: deployment. Observable result: single-user operation with scheduled jobs. Failure/recovery: missing key skips the service with a log. Continuation: additive product mode later.

FR-10 — Key handling (explicit) As the Operator / Reviewer, I should keep keys in backend .env only — never in frontend, logs, fixtures, or git — mask them in logs, and ship .env.example, so that credentials stay safe.

  • Keys: INDIANAPI_JOBS_KEY, TYPESAFE_API_KEY, LLM_API_KEY with LLM_PROVIDER=openai|anthropic, optional YOUTUBE_API_KEY, GITHUB_TOKEN, and alert credentials.
  • Trigger: service start. Observable result: masked logs; .env.example present. Failure/recovery: missing key skips the service with a log. Continuation: service resumes when the key is provided.

FR-11 — IndianAPI job data (explicit) As the Job Seeker, I should have job data fetched from GET https://jobs.indianapi.in/jobs with header X-Api-Key, so that postings are current and sourced.

  • Confirmed param: limit, sent as a STRING (e.g. "50"); an int gives 422. Unconfirmed: title, location, company, experience, job_type, pagination. Response array fields: id, title, company, about_company, job_description, job_title, job_type, location, experience, role_and_responsibility, education_and_skills, apply_link, posted_date (ISO 8601). No salary field: show "Not available from source", never estimate. Rate limits unknown: limiter default 1 req/2s, INDIANAPI_DAILY_CAP, retry 429/5xx max 3 with backoff+jitter, honor Retry-After.
  • Trigger: scheduled poll. Observable result: normalized jobs stored. Failure/recovery: retries with backoff; failures logged. Continuation: next scheduled run.

FR-12 — TypeSafe Jev decisions (explicit) As the Job Seeker, I should have structured decisions produced by TypeSafe Jev at POST https://api.typesafe.ai/v1/systemone with Bearer auth, so that enrichment and fit use typed decisions rather than free text.

  • Use typesafe-sdk (AsyncTypeSafeClient, Choice/Score/Noul, RetryPolicy). Body {state, model, questions}. choice → choice + probabilities + confidence; score → score + probabilities + confidence; noul → probability 0–1. Pin JEV_MODEL to an exact id via SDK list-models (not jev-latest). Store model version per record. First-install skill: claude plugin marketplace add typesafe-ai/skills && claude plugin install typesafe@typesafe-ai. Read docs.typesafe.ai primitives, confidence, patterns (fan-out, confidence routing, composite scoring), cookbooks (rerank, entity alignment, pre-parsed extraction), and jev-1.13 jaggedness; summarize limits in NOTES.md.
  • Trigger: enrichment or fit run. Observable result: typed answers with confidence and model version. Failure/recovery: API failure → status failed, retry next run, never block ingestion. Continuation: retry next run.

FR-13 — LLM text (explicit) As the Job Seeker, I should have text generated by a provider-agnostic LLM at temperature 0.2, JSON-schema-validated and cached, so that plans and questions are consistent and cheap to regenerate.

  • Trigger: plan or question generation. Observable result: validated, cached text. Failure/recovery: validation failure rejects the output. Continuation: retry.

FR-14 — Prep resources (explicit) As the Job Seeker, I should have prep resources collected from YouTube Data API, Dev.to, HN Algolia, arXiv, GitHub search, and RSS, storing title, URL, date, metadata, and a short summary in our own words, so that prep is grounded and legally safe.

  • Never store full copyrighted text. Trigger: prep run. Observable result: resource rows with metadata and summaries. Failure/recovery: provider failure skips that provider with a log. Continuation: next prep run.

FR-15 — Hard rules (explicit) As the Operator / Reviewer, I should have the hard rules enforced: no scraping any website (no LinkedIn/Indeed/Naukri/Glassdoor); missing field = NULL; no fake data outside fixtures; every job shows source + apply_link; low-confidence decisions → "needs review", never silently accepted; market pages show the banner "Source: IndianAPI only; trends reflect this source, not the full market."

  • Trigger: every pipeline and render. Observable result: compliant data and UI. Failure/recovery: violations are treated as defects. Continuation: enforced continuously.

FR-16 — Stack (explicit) As the Operator / Reviewer, I should have the specified stack used: Python 3.11, FastAPI, SQLAlchemy 2, Alembic, httpx, Pydantic v2, APScheduler, PostgreSQL 16 + pgvector, rapidfuzz, typesafe-sdk; Next.js 14, TypeScript, Tailwind, TanStack Query, lightweight-charts, Recharts; Docker Compose, pytest, ruff, mypy, GitHub Actions.

  • Trigger: build. Observable result: stack in use. Failure/recovery: deviations require asking first. Continuation: maintained across phases.

FR-17 — Config (explicit) As the Operator / Reviewer, I should have config/*.yaml hold profile (skills {name, level 1-5}, years, domains, cities, remote pref, target roles, min seniority, deal-breakers, resume_text), roles (canonical + description + synonyms + regex), skills (id, name, description), locations (aliases, e.g. Bengaluru → Bangalore), thresholds, fit_weights, schedule, and sources, so that behavior is configurable without code changes.

  • Trigger: startup and profile edits. Observable result: config loaded and applied. Failure/recovery: invalid config fails fast with a clear error. Continuation: reload on change.

FR-18 — Phase 0 probes (explicit) As the Operator / Reviewer, I should have probe_indianapi.py run limit="5", then test each unconfirmed param and pagination, recording status, count, whether results truly filtered, posted_date freshness, id stability, and max limit, saving to tests/fixtures/indianapi/; and probe_jev.py run 10 fixture JDs through the enrichment questions, measuring latency, tokens, confidence spread, batched vs split calls, and state-length limit, saving to tests/fixtures/jev/; with NOTES.md recording results, query plan, and batch size.

  • Trigger: Phase 0 start. Observable result: fixtures and NOTES.md entries. Failure/recovery: probe failures are reported, not guessed. Continuation: stop and report.

FR-19 — Pipeline (explicit) As the Job Seeker, I should have the pipeline run fetch → normalize → loose regex pre-filter → dedup → Jev enrich → gate → upsert → embed → fit → metrics → alerts, so that jobs flow from source to ranked, tracked state.

  • Query plan: if filters work, loop roles × cities within cap; else fetch max and filter locally. Normalize: source_job_id=id; title=job_title or title; keep JD sections separate; location → city/state/country; posted_at=parse(posted_date); raw JSON in jobs.raw. Dedup: sha256(norm company+title+city); rapidfuzz title ≥85, same company, 30 days → Jev same-posting score 0/1/2: 2 merge (keep earliest posted_at, all links), 1 review. Unseen in 3 full syncs → inactive.
  • Trigger: scheduled run. Observable result: normalized, deduplicated, enriched, ranked jobs. Failure/recovery: per-stage failures logged and retried next run. Continuation: next scheduled run.

FR-20 — Jev enrichment (explicit) As the Job Seeker, I should have each job enriched with batched Jev questions using state = title, company, location, experience, job_type, JD sections (truncated per Phase 0), so that structured attributes are attached.

  • Questions: is_ai_role (noul) — primary work is AI/ML/DS/LLM/CV/NLP/MLOps engineering; role (choice) — roles.yaml + other; seniority (score) — fresher, junior 0-2y, mid 2-5y, senior 5-8y, lead 8y+; work_mode (choice) — onsite|hybrid|remote|unspecified; skill_<id> (noul each, chunked) — job requires/prefers <skill>; experience (choice) — regex-extracted spans + none, code converts to min/max; domain (choice) — fintech, real estate, hospitality, healthcare, e-commerce, HR tech, document AI, manufacturing, SaaS, other.
  • Gating: is_ai_role ≥0.7 keep, 0.4–0.7 review, <0.4 drop (logged). Skill ≥0.6 attach with prob. Low confidence → NULL + needs_review (or LLM fallback if enabled). Cache by sha256(state+question_set_version+model). API failure → status failed, retry next run, never block ingestion.
  • Trigger: enrichment run. Observable result: enriched job rows with confidences. Failure/recovery: as above. Continuation: retry next run.

FR-21 — Fit scoring (explicit) As the Job Seeker, I should have one Jev call per job with state = JD + profile summary producing skill_match score 0–3, experience_match score 0–2 (below/match/above), domain_match noul, growth_potential score 0–2 (scope beyond my level), and red_flags noul (unrealistic/vague/mislabeled); with code applying location/remote prefs, deal-breakers (hard filter), and min seniority; fit 0–100 = weighted sum (fit_weights); components and confidence stored; UI shows "why this score"; recompute on profile change.

  • Trigger: fit run or profile change. Observable result: fit score with components and confidence. Failure/recovery: missing components flagged. Continuation: recompute on profile change.

FR-22 — Skill gap (explicit) As the Job Seeker, I should have per-skill 30/90-day trend, senior-vs-mid share, share in my top-fit jobs, and my level; gap = high on these AND my level ≤2; /growth shows top 5 gaps, ranked resources, and a 4-week plan (LLM, grounded only in matched resources, cite URLs).

  • Trigger: gap run. Observable result: gap rows and plan. Failure/recovery: insufficient history labeled. Continuation: daily updates.

FR-23 — Prep per job (explicit) As the Job Seeker, I should have a shortlist of top 30 by 0.5*cosine + 0.5*skill overlap, a Jev rerank score per (JD, resource) pair (not/somewhat/very useful), and the top 10 kept; POST /jobs/{id}/interview-questions uses the LLM to produce 15 questions (technical, ML system design, behavioral), tagged with my gap skills, cached.

  • Trigger: prep run or question request. Observable result: ranked prep and cached questions. Failure/recovery: generation failure shows inline error. Continuation: retry.

FR-24 — Tracker (explicit) As the Job Seeker, I should have statuses saved, applied, interview, offer, rejected, withdrawn with dates, notes, and next action, plus a funnel chart.

  • Trigger: tracking action. Observable result: application row and funnel update. Failure/recovery: save failure preserves edits. Continuation: status changes over time.

FR-25 — Alerts (explicit) As the Job Seeker, I should receive a daily digest (email|telegram) of new jobs with fit ≥ ALERT_MIN_FIT plus the top 3 rising gap skills.

  • Trigger: 08:30 Asia/Kolkata schedule. Observable result: digest delivered through the configured channel. Failure/recovery: delivery failure logged; retry next run. Continuation: daily.

FR-26 — Market metrics (explicit) As the Job Seeker, I should see per skill/role/city: active, new today, delta% vs yesterday and 7-day average; AI Job Index = active AI jobs, base 100 day 1; volume swing >50% → "source anomaly", excluded from deltas; Jev model change → chart marker; "Insufficient history" until 7 days.

  • Trigger: metrics run. Observable result: metrics rows and chart. Failure/recovery: anomalies excluded; insufficient history labeled. Continuation: daily.

FR-27 — Schedule (explicit) As the Operator / Reviewer, I should have the schedule run in Asia/Kolkata: poll every 2h 08:00–22:00; full sync 02:00; prep 03:00; metrics+fit+gap 04:00 and after runs; digest 08:30; with ingestion_runs and api_usage (service, calls, tokens) logged.

  • Trigger: clock. Observable result: scheduled runs and logs. Failure/recovery: missed runs logged. Continuation: next scheduled run.

FR-28 — Data model (explicit) As the Operator / Reviewer, I should have the specified tables: companies, jobs (role, seniority, work_mode, domain, ai_role_prob, exp_min/max, confs, needs_review, model_version, raw, embedding, is_active), job_sources, skills, job_skills(prob), job_fit(components JSONB), skill_gap_daily, prep_resources, resource_skills, job_prep_rank, interview_question_sets, applications, review_queue, daily_metrics, ingestion_runs, api_usage, decision_cache, alerts_sent; unique (source, source_job_id); indexes on posted_at, role, city, is_active, dedup_hash, fit; pgvector.

  • Trigger: migrations. Observable result: schema present. Failure/recovery: migration failure reported. Continuation: next migration.

FR-29 — API (explicit) As the Job Seeker, I should have /api/v1 endpoints: GET /jobs (filters role, skill, city, work_mode, seniority, min_fit, posted_within_days, q; sort fit|date), GET /jobs/{id}, GET /jobs/{id}/prep, POST /jobs/{id}/interview-questions, GET/PUT /profile, GET /growth/gaps, GET /growth/plan/{skill}, GET/POST/PATCH /applications, GET /market/{index,ticker,skills,locations}, GET /prep, GET /admin/{health,runs,api-usage,review-queue}, POST /admin/review/{id} (saved as labels).

  • Trigger: client requests. Observable result: typed responses. Failure/recovery: standard error responses. Continuation: normal operation.

FR-30 — Frontend (explicit) As the Job Seeker, I should have the routes / (my top-fit new jobs + gaps), /jobs, /jobs/[id] (JD, fit breakdown, matched/missing skills, apply, prep, questions, track), /growth, /applications, /market (ticker tape, index chart, gainers/decliners, banner), /skills/[name], /prep, /profile, /admin; responsive, dark mode, with loading/empty/error states.

  • Trigger: navigation. Observable result: rendered pages. Failure/recovery: error states with retry. Continuation: normal navigation.

FR-31 — Tests (explicit) As the Operator / Reviewer, I should have tests using fixtures only with no live calls in CI, covering normalize, string limit, error codes, Jev parsing/gating/cache/fallback, dedup, experience regex, fit, gap, metrics, anomaly guard, and alerts; an eval script where I label 50 JDs (is_ai_role, role, skills, fit 1-5) reporting precision/recall per threshold and fit correlation to tune thresholds and weights; coverage ≥80%.

  • Trigger: CI. Observable result: test and eval reports. Failure/recovery: failures reported as open issues. Continuation: next phase.

FR-32 — Output rules (explicit) As the Operator / Reviewer, I should have files listed before each phase; a missing key skips the service with a log; and any new service or dependency requires asking first.

  • Trigger: phase start or dependency need. Observable result: file list, skip log, or question. Failure/recovery: unapproved additions are not made. Continuation: proceed after approval.

FR-33 — Self-service enrollment (required_inference) As the Job Seeker, I should be able to establish my own identity through self-service enrollment, so that my profile, fit history, applications, and review labels remain bound to me and resumable.

  • Trigger: first use. Observable result: identity created. Failure/recovery: validation errors shown inline. Continuation: proceed to profile setup.

FR-34 — Returning verification (required_inference) As the Job Seeker, I should verify my identity on return before accessing protected profile, job, growth, preparation, application, market, or administrative work, so that my durable state stays private and continuous.

  • Trigger: access to a protected destination. Observable result: session established. Failure/recovery: invalid credentials shown inline. Continuation: proceed to the intended destination.

FR-35 — Operator authorization (required_inference) As the Operator / Reviewer, I should be authorized before accessing review queues, operational health, ingestion runs, API usage, and phase probe reporting, so that operational functions remain under my control.

  • Trigger: access to /admin. Observable result: authorized access. Failure/recovery: unauthorized access denied. Continuation: normal operation.

FR-36 — Digest delivery configuration (required_inference) As the Job Seeker, I should have external email or Telegram delivery configured for the daily digest, with delivery remaining provider/external-channel owned.

  • Trigger: digest schedule. Observable result: digest delivered. Failure/recovery: delivery failure logged. Continuation: daily.
Page 10 of 23

4. User Personas

Page 11 of 23

Job Seeker (AI/ML candidate)

Product context. The single personal-mode user of AI Career Pulse. They are an AI/ML candidate who reads tables, scores, and charts all day and will instantly notice a template. They operate the tool for their own job search, not for a demo.

Primary goal. A shortlist of well-matched roles with grounded prep material and an up-to-date application pipeline.

Distinct accepted responsibilities. Maintaining the profile (skills with levels 1–5, years, domains, cities, remote preference, target roles, minimum seniority, deal-breakers, resume text); reviewing fit breakdowns and "why this score" explanations; resolving low-confidence/needs-review items; acting on the daily digest of new high-fit jobs and rising gap skills; preparing for specific job interviews; and tracking applications through the funnel.

Relevant inputs or decisions. Profile fields and skill levels; filter and sort choices on /jobs; decisions to apply, save, or track; decisions to generate interview questions; decisions to open ranked resources and follow the 4-week plan.

Interactions with other accepted participants. The Job Seeker is the sole human participant in the career workflows. Their work is supported by external providers (IndianAPI for job data, TypeSafe Jev for decisions, the LLM for text, prep-resource providers) and by the internal pipeline. The Operator / Reviewer is the same person acting operationally.

Observable success. A ranked job list with transparent fit scores; a top-5 gap list with a grounded 4-week plan; cached interview questions tagged with gap skills; an application funnel that reflects current statuses; and a market dashboard read with its source limits clearly in view.

What makes this role different. The Job Seeker's work is analytical and personal: every number on screen is about them — their level, their fit, their gaps, their pipeline. The interface is an instrument for one person, and the value comes from the transparency of the score and the honesty of the source limits.

Page 12 of 23

Operator / Reviewer (admin)

Product context. The same personal-mode owner acting in an operational capacity. They run the Phase 0 probes, read NOTES.md unknowns and assumptions, and review each phase's test results, summary, and open issues before the build continues.

Primary goal. A healthy, observable pipeline with no silently accepted low-confidence decisions and no invented endpoints, parameters, fields, or data.

Distinct accepted responsibilities. Monitoring ingestion runs, API usage, health, and the review queue; labeling review-queue items via POST /admin/review/{id} so labels feed threshold and weight tuning; running the Phase 0 probes; reading NOTES.md; and reviewing each phase's test results, summary, and open issues.

Relevant inputs or decisions. Review-queue items and their labels; ingestion-run and API-usage records; probe results; phase test summaries and open issues; decisions to approve or reject new services or dependencies.

Interactions with other accepted participants. The Operator / Reviewer is the same person as the Job Seeker, acting operationally. Their work is supported by the internal pipeline and by the external providers whose usage they monitor.

Observable success. A healthy pipeline with visible runs and API usage; a review queue that is labeled rather than ignored; probe results and NOTES.md entries that record unknowns honestly; and phase reviews that stop the build until issues are addressed.

What makes this role different. The Operator / Reviewer's work is about the system rather than the job search: they watch the machine, label its uncertain outputs, and decide whether the build may continue. Their success is measured in observability and honesty, not in fit scores.

Page 13 of 23

5. Core User Flows

Flow 1 — First use: enroll and set up the profile (Job Seeker)

  1. The Job Seeker opens / and reads the hero band: the headline, the live data line, and the top-three fit rows (or a neutral "no data yet" line if nothing has been ingested).
  2. They choose to enter the product and land on Sign Up.
  3. They complete self-service enrollment. On success, they proceed to profile setup.
  4. On /profile, they enter skills with levels 1–5, years, domains, cities, remote preference, target roles, minimum seniority, deal-breakers, and resume text, and save.
  5. Saving triggers recomputation of fit on profile change. The score-composition schematic shows how the score is composed from skill_match, experience_match, domain_match, growth_potential, and red_flags.
  6. Failure/recovery: validation or save failure shows an inline message and preserves unsaved changes; they correct and resubmit.
  7. Continuation: they proceed to /jobs to review ranked roles.

Flow 2 — Returning use: verify and resume (Job Seeker)

  1. The Job Seeker returns and opens a protected destination.
  2. They land on Login and submit credentials.
  3. On success, they are taken to the intended destination with their durable state intact.
  4. Failure/recovery: invalid credentials show an inline message with no field echo; they retry or navigate to Sign Up.
  5. Continuation: normal work resumes.
Page 14 of 23

Flow 3 — Rank and review jobs by fit (Job Seeker)

  1. The Job Seeker opens /jobs.
  2. They apply filters (role, skill, city, work_mode, seniority, min_fit, posted_within_days, q) and choose a sort (fit or date).
  3. The ruled job table shows company, role, city, posted date, source, apply link, and the fit numeral right-aligned in a fixed-width column with a 3px probability bar beneath it.
  4. They open a job to see the JD and fit breakdown on the left at 62% and matched/missing skills, prep resources, and tracker controls in the sticky right column at 38%.
  5. The fit breakdown shows components and confidence, and the "why this score" explanation. The salary line reads "Not available from source".
  6. Failure/recovery: fetch failure shows an inline error with retry; empty results show active filters and a clear-filters action.
  7. Continuation: they open the apply link or track the application.

Flow 4 — Read skill gaps and follow a learning plan (Job Seeker)

  1. The Job Seeker opens /growth.
  2. They read the continuous ruled list of gap rows: skill name, a bar split into a tangerine filled portion (my level 1–5) and an acid-lime demand portion (30-day market share), and the numeric gap right-aligned in tabular figures.
  3. They review the top 5 gaps, the ranked resources, and the 4-week plan grounded only in matched resources with cited URLs.
  4. They open a skill detail on /skills/[name] to see the 30/90-day trend, senior-versus-mid share, share in their top-fit jobs, and their level.
  5. Failure/recovery: insufficient history shows "Insufficient history" until 7 days; plan generation failure shows resources without a plan and an inline error.
  6. Continuation: they open a ranked resource or return to /growth.

Flow 5 — Prepare for a specific job (Job Seeker)

  1. From /jobs/[id], the Job Seeker reviews the JD and fit breakdown.
  2. They open prep resources in the sticky right rail, which shows the shortlist of top 30 by 0.5*cosine + 0.5*skill overlap and the Jev-reranked top 10 with rerank scores (not/somewhat/very useful).
  3. They request interview questions. The LLM produces 15 questions (technical, ML system design, behavioral) tagged with their gap skills, cached.
  4. Failure/recovery: generation failure shows an inline error and preserves the apply link; they retry.
  5. Continuation: they track the application or return to /jobs.
Page 15 of 23

Flow 6 — Track an application (Job Seeker)

  1. From /jobs/[id], the Job Seeker saves or tracks the application.
  2. On /applications, they see the application row with status (saved, applied, interview, offer, rejected, withdrawn), dates, notes, and next action.
  3. They change status and edit dates, notes, and next action as the process advances.
  4. The funnel chart updates to reflect current statuses.
  5. Failure/recovery: save failure shows an inline error and preserves edits; they retry.
  6. Continuation: they continue updating as the process advances.

Flow 7 — Read the market dashboard (Job Seeker)

  1. The Job Seeker opens /market.
  2. They read the full-bleed ticker tape (pausing on hover), the AI Job Index chart (base 100 on day 1), and the gainers/decliners table.
  3. The permanent source-limits banner reads "Source: IndianAPI only; trends reflect this source, not the full market."
  4. Volume swings greater than 50% are marked "source anomaly" and excluded from deltas; Jev model changes are marked on the chart; "Insufficient history" appears until 7 days.
  5. Failure/recovery: metrics fetch failure shows an inline error while the banner remains visible; they retry on the next metrics run.
  6. Continuation: they open a skill or location context.

Flow 8 — Receive and act on the daily digest (Job Seeker)

  1. At 08:30 Asia/Kolkata, the digest is delivered through the configured email or Telegram channel.
  2. The digest lists new jobs with fit ≥ ALERT_MIN_FIT plus the top 3 rising gap skills.
  3. The Job Seeker opens a listed job on /jobs/[id] and reviews the fit breakdown.
  4. Failure/recovery: delivery failure is logged and retried on the next run.
  5. Continuation: they track the application or open /growth to act on a rising gap.
Page 16 of 23

Flow 9 — Monitor the pipeline and label the review queue (Operator / Reviewer)

  1. The Operator / Reviewer opens /admin and reviews health, ingestion runs, API usage (service, calls, tokens), and the review queue.
  2. They label a review-queue item via POST /admin/review/{id}; labels are saved and feed threshold and weight tuning.
  3. They read the Phase 0 probe reporting surfaced on /admin.
  4. Failure/recovery: fetch failure shows an inline error with retry.
  5. Continuation: they return to monitoring or proceed to the next phase review.

Flow 10 — Run Phase 0 probes and review the phase (Operator / Reviewer)

  1. The Operator / Reviewer runs probe_indianapi.py with limit="5", then tests each unconfirmed param and pagination, recording status, count, whether results truly filtered, posted_date freshness, id stability, and max limit, saving to tests/fixtures/indianapi/.
  2. They run probe_jev.py with 10 fixture JDs through the enrichment questions, measuring latency, tokens, confidence spread, batched vs split calls, and state-length limit, saving to tests/fixtures/jev/.
  3. They record results, query plan, and batch size in NOTES.md.
  4. After each phase, tests run, results are summarized, open issues are listed, and the build stops for their review.
  5. Failure/recovery: probe failures are reported, not guessed; unknowns are logged in NOTES.md and raised as questions.
  6. Continuation: the next phase begins only after review.
Page 17 of 23

6. Visuals Colors and Theme

The creative direction is authoritative for this section. The muse is Rasmus Andersson: systematic product craft with an opinion — tight 4/8-pt scale, tabular numerals as ornament, hairline rules, keyboard-first density, and a palette that deliberately refuses bootstrap blue on white. The headline register is serious craft, quiet confidence, instrument-grade precision, zero marketing gloss.

Mode. Dark mode is the primary mode.

Colour tokens by role (dark mode).

RoleTokenHex
Background--bg#141517
Surface (panels, cards, ticker tape)--surface#1C1E21
Text (body, numerals)--text#F2F0EC
Primary action (apply, active nav, index line, selected chip)--primary#E8552B
Data signal (gainers, rising gap bars, fit ≥80, delta up)--accent#C6F24E
Muted (labels, metadata, "Not available from source", timestamps, axis text)--muted#8A8D93
Hairline rules--rule#2A2D31
Decliners and low-confidence flags--danger#FF5C5C

Tangerine #E8552B is the single primary action colour, used on roughly 5% of pixels and never as a fill for large blocks. Acid lime #C6F24E is the data signal. Muted #8A8D93 carries labels and metadata. Hairlines are #2A2D31 at 1px. Red #FF5C5C is used only inside charts and the needs-review pill. No blue anywhere in the system.

Typography.

  • Headings: Inter Tight, 600–700 weight, tight tracking (−0.02em to −0.04em), sentence case for section headings and uppercase for micro-labels at 11px with +0.08em tracking.
  • Hero headline: oversized, clamp(3.5rem, 9vw, 8.5rem), flush left, ragged right, with the fit score numeral interleaved at the same optical weight so type and data are one gesture.
  • Body: Instrument Sans.
  • Numerals: tabular figures via font-feature-settings: 'tnum' 1, 'ss01' 1.
  • Monospace numerals: JetBrains Mono for ticker tape, index values, probabilities, and API latency in /admin.
  • Type scale: 1.25 modular with a data tier — 96 / 64 / 40 / 28 / 20 / 16 / 14 / 12, plus an 11px uppercase micro-label tier.
  • Line-height: 1.05 display, 1.25 headings, 1.55 body, 1.15 table rows.

Shape language. Sharp and exact. 6px radius on buttons and inputs, 10px on panels, 0px on the ticker tape, table rows, and chart frames. Hairline 1px borders (#2A2D31) do all the separation — no shadows, no glow, no blur. Focus rings are a 2px #C6F24E outline offset by 2px, keyboard-first. Confidence is shown as a thin horizontal probability bar with a hard right edge, never as a pill or a badge.

Spacing rhythm. 4/8-pt scale; 12-column grid, 24px gutters, 1440px max content, with charts and the ticker tape full-bleed edge to edge. Left rail navigation, 240px, persistent, with section labels in 11px uppercase and the current page marked by a 2px tangerine left rule.

Imagery style. No photography and no illustration. The imagery is the interface itself: the fit-score numeral as the hero object, ruled tables, monospace ticker tape, hairline chart grids, a small SVG sparkline per skill row, and one schematic diagram on /profile showing how the score is composed (weighted bars labelled skill_match, experience_match, domain_match, growth_potential, red_flags). Icons are 16px 1.5px-stroke line icons. Every chart is labelled with its source and its limits inline — the market banner is part of the visual system, not a dismissible toast.

Page 18 of 23

7. Signature Design Concept

The public entry (/) is a full-bleed graphite band, edge to edge, with a 1px hairline top and bottom. It is not a centred headline with a button.

  • Left side: a flush-left Inter Tight headline at clamp(3.5rem, 9vw, 8.5rem) reading "YOUR NEXT ROLE, SCORED." with the word SCORED in #E8552B. Immediately under it, a single ruled line of live data in JetBrains Mono at 14px: AI JOB INDEX 100.0 · ACTIVE 1,284 · NEW TODAY 37 · +2.1% vs 7D, with the delta in #C6F24E.
  • Right side: the top three fit-ranked jobs as three hairline-ruled rows — company, role, city on the left; an oversized tabular fit numeral (e.g. 91, 88, 84) right-aligned in #C6F24E with a 3px probability bar beneath it.
  • No gradient, no blob, no stock photo, no illustration — the composition is type plus numerals plus rules on graphite.

The concept recomposes only accepted content, states, and controls: the headline, the live data line drawn from daily_metrics, and the top-three fit rows drawn from job_fit. It introduces no new behavior, page, or destination.

Page 19 of 23

8. Interaction Model & Motion Direction

Interaction Model: Static Motion Tempo: restrained Hero Dimensionality: flat

Landing Hero Motion Brief. The focal subject is the fit-score numeral and the ruled data line, not an illustration. The input→transformation→outcome thesis: on first paint, the fit numeral counts up once over 400ms with no bounce, and the ruled data line settles in; the outcome is a composed first frame where type and data read as one gesture. The motion vocabulary is restrained and functional — 120–200ms ease-out on everything, filter chips settle in, table rows fade in on data load, chart lines draw in left-to-right once, and the ticker tape scrolls continuously at a slow constant rate and pauses on hover. The only expressive moment is the index line drawing itself on first load of /market. There is no parallax, no hover-lift, no spring, and no scroll-triggered choreography. The reduced-motion state removes the count-up and the ticker scroll, rendering the final numerals and a static ticker line instead.

Page 20 of 23

9. Non-Functional Requirements

  • NFR-1 — No scraping (explicit). No scraping of any website, including LinkedIn, Indeed, Naukri, and Glassdoor. Rationale: legal and ethical boundary set by the owner.
  • NFR-2 — Missing field = NULL (explicit). Missing fields are stored as NULL; no fake data exists outside fixtures. Rationale: data honesty.
  • NFR-3 — Source and apply link (explicit). Every job shows its source and apply_link. Rationale: traceability.
  • NFR-4 — Low-confidence handling (explicit). Low-confidence decisions become "needs review" and are never silently accepted. Rationale: correctness of downstream ranking.
  • NFR-5 — Market banner (explicit). Market pages show "Source: IndianAPI only; trends reflect this source, not the full market." Rationale: source-limit honesty.
  • NFR-6 — No salary estimation (explicit). No salary field exists in the source; show "Not available from source" and never estimate. Rationale: no invented data.
  • NFR-7 — String limit (explicit). The IndianAPI limit parameter must be sent as a STRING; an int gives 422. Rationale: confirmed API behavior.
  • NFR-8 — Jev decisions only (explicit). Jev generates no text; pin JEV_MODEL to an exact id via SDK list-models (not jev-latest); store model version per record. Rationale: reproducibility and correct role separation.
  • NFR-9 — LLM text only (explicit). The LLM is text-only, provider-agnostic, temperature 0.2, JSON-schema-validated, cached. Rationale: consistency and cost control.
  • NFR-10 — Prep resource storage (explicit). Store title, URL, date, metadata, and a short summary in our own words; never full copyrighted text. Rationale: legal safety.
  • NFR-11 — No invented endpoints/params/fields/data (explicit). Log unknowns and assumptions in NOTES.md and ask. Rationale: correctness.
  • NFR-12 — Phase review stops (explicit). After each phase: run tests, summarize, list open issues, stop for review. Rationale: controlled delivery.
  • NFR-13 — Output rules (explicit). List files before each phase; a missing key skips the service with a log; ask before adding any new service or dependency. Rationale: controlled scope.
  • NFR-14 — Tests (explicit). Fixtures only, no live calls in CI; coverage ≥80%. Rationale: deterministic CI.
  • NFR-15 — Gating thresholds (explicit). is_ai_role ≥0.7 keep, 0.4–0.7 review, <0.4 drop (logged); skill ≥0.6 attach with prob. Rationale: precision control.
  • NFR-16 — Dedup (explicit). rapidfuzz title ≥85 with same company within 30 days → Jev same-posting score 0/1/2 (2 merge keeping earliest posted_at and all links, 1 review); unseen in 3 full syncs → inactive. Rationale: duplicate control.
  • NFR-17 — Market anomaly guard (explicit). Volume swing >50% → "source anomaly", excluded from deltas; "Insufficient history" until 7 days. Rationale: trend honesty.
  • NFR-18 — IndianAPI rate handling (explicit). Rate limits unknown: limiter default 1 req/2s, INDIANAPI_DAILY_CAP, retry 429/5xx max 3 with backoff+jitter, honor Retry-After. Rationale: safe consumption.
  • NFR-19 — Failure isolation (explicit). API failure → status failed, retry next run, never block ingestion. Rationale: pipeline resilience.
  • NFR-20 — Key handling (explicit). Keys live in backend .env only; never in frontend, logs, fixtures, or git; mask in logs; ship .env.example. Rationale: credential safety.
  • NFR-21 — Personal mode (explicit). Single user, APScheduler, PostgreSQL; no Celery, no Redis, no SSE; interfaces stay clean so a later product mode is additive. Rationale: simplicity with a growth path.
  • NFR-22 — Responsive dark mode (explicit). The frontend is responsive, dark mode, with loading/empty/error states. Rationale: usability.
  • NFR-23 — Accessibility of the palette (required_inference). Body text #F2F0EC on #141517 yields roughly 13:1 contrast, safe at 14px; focus rings are a 2px #C6F24E outline offset by 2px. Rationale: readable dense data surfaces.
Page 21 of 23

10. Tech Stack

Source-specified choices are preserved exactly.

Backend. Python 3.11, FastAPI, SQLAlchemy 2, Alembic, httpx, Pydantic v2, APScheduler, PostgreSQL 16 + pgvector, rapidfuzz, typesafe-sdk.

Frontend. Next.js 14, TypeScript, Tailwind, TanStack Query, lightweight-charts, Recharts.

Tooling and delivery. Docker Compose, pytest, ruff, mypy, GitHub Actions.

External services. IndianAPI jobs endpoint (GET https://jobs.indianapi.in/jobs, header X-Api-Key); TypeSafe Jev (POST https://api.typesafe.ai/v1/systemone, Bearer auth); a provider-agnostic LLM (LLM_PROVIDER=openai|anthropic); prep-resource providers (YouTube Data API, Dev.to, HN Algolia, arXiv, GitHub search, RSS); and an email or Telegram channel for the daily digest.

Configuration. config/*.yaml for profile, roles, skills, locations, thresholds, fit_weights, schedule, and sources.

Page 22 of 23

11. Assumptions and Constraints

Assumptions.

  • A-1 (explicit). The IndianAPI limit parameter is sent as a STRING; an int gives 422. All other parameters (title, location, company, experience, job_type, pagination) are unconfirmed until Phase 0 probes report.
  • A-2 (explicit). IndianAPI rate limits are unknown; the limiter defaults to 1 req/2s with INDIANAPI_DAILY_CAP and bounded retries.
  • A-3 (explicit). JEV_MODEL is pinned to an exact id obtained via SDK list-models, never jev-latest.
  • A-4 (explicit). The LLM provider is selected by LLM_PROVIDER=openai|anthropic; text generation is temperature 0.2, JSON-schema-validated, and cached.
  • A-5 (explicit). Prep-resource providers are used only for metadata and short summaries in our own words.
  • A-6 (required_inference). Self-service enrollment is the truthful first-use path because no invitation or provisioning boundary is established.
  • A-7 (required_inference). The daily digest requires external email or Telegram delivery configuration; delivery remains provider/external-channel owned.
  • A-8 (required_inference). Operator authorization gates review queues, health, runs, API usage, and phase probe reporting.

Constraints.

  • C-1 (explicit). Personal mode is the default: single user, APScheduler, PostgreSQL; no Celery, no Redis, no SSE. Interfaces must stay clean so a later product mode (auth, Celery, SSE) is additive.
  • C-2 (explicit). Keys live in backend .env only; never in frontend, logs, fixtures, or git; mask keys in logs; ship .env.example.
  • C-3 (explicit). No scraping any website (no LinkedIn/Indeed/Naukri/Glassdoor).
  • C-4 (explicit). Missing field = NULL; no fake data outside fixtures.
  • C-5 (explicit). Every job shows source + apply_link.
  • C-6 (explicit). Low-confidence decisions → "needs review", never silently accepted.
  • C-7 (explicit). Market pages show the banner "Source: IndianAPI only; trends reflect this source, not the full market."
  • C-8 (explicit). No salary field from IndianAPI: show "Not available from source", never estimate.
  • C-9 (explicit). The IndianAPI limit param must be sent as a STRING (int gives 422).
  • C-10 (explicit). Jev is decisions only and generates NO text; pin JEV_MODEL to an exact id via SDK list-models (not jev-latest); store model version per record.
  • C-11 (explicit). The LLM is text only, provider-agnostic, temperature 0.2, JSON-schema-validated, cached.
  • C-12 (explicit). Prep resources: store title, URL, date, metadata, and a short summary in our own words; never full copyrighted text.
  • C-13 (explicit). Never invent endpoints, params, fields, or data; log unknowns/assumptions in NOTES.md and ask the user.
  • C-14 (explicit). After each phase: run tests, summarize, list open issues, stop for the user's review.
  • C-15 (explicit). List files before each phase; missing key → skip service with log; ask before adding any new service or dependency.
  • C-16 (explicit). Tests use fixtures only, no live calls in CI; coverage ≥80%.
  • C-17 (explicit). Gating thresholds: is_ai_role ≥0.7 keep, 0.4–0.7 review, <0.4 drop (logged); skill ≥0.6 attach with prob.
  • C-18 (explicit). Dedup: rapidfuzz title ≥85 with same company within 30 days → Jev same-posting score 0/1/2 (2 merge keeping earliest posted_at and all links, 1 review); unseen in 3 full syncs → inactive.
  • C-19 (explicit). Market volume swing >50% → "source anomaly", excluded from deltas; "Insufficient history" until 7 days.
  • C-20 (explicit). IndianAPI rate limits unknown: limiter default 1 req/2s, INDIANAPI_DAILY_CAP, retry 429/5xx max 3 with backoff+jitter, honor Retry-After.
  • C-21 (explicit). API failure → status failed, retry next run, never block ingestion.
  • C-22 (explicit). The generic indigo/blue-on-white SaaS template is forbidden for this project; no blue or indigo appears anywhere in the accent system, including chart series and links.
Page 23 of 23

12. Glossary

  • AI Career Pulse — the personal, single-user AI/ML job intelligence application described in this document.
  • Personal mode — the default operating mode: one user, APScheduler, PostgreSQL, no Celery/Redis/SSE.
  • IndianAPI — the job data source at GET https://jobs.indianapi.in/jobs, authenticated with header X-Api-Key.
  • TypeSafe Jev — the decisions-only structured decision service at POST https://api.typesafe.ai/v1/systemone, used with typesafe-sdk.
  • JEV_MODEL — the pinned exact Jev model id obtained via SDK list-models; never jev-latest.
  • Noul / Choice / Score — the three Jev question primitives: noul returns a probability 0–1; choice returns a choice with probabilities and confidence; score returns a score with probabilities and confidence.
  • Fit score — the 0–100 weighted sum of skill_match (0–3), experience_match (0–2), domain_match (noul), growth_potential (0–2), and red_flags (noul), combined with code-applied location/remote preferences, deal-breakers, and minimum seniority.
  • fit_weights — the configured weights used in the fit weighted sum.
  • Needs review — the state assigned to low-confidence decisions; never silently accepted.
  • Review queue — the list of items awaiting operator labeling via POST /admin/review/{id}.
  • Gating — the threshold rules applied to is_ai_role (≥0.7 keep, 0.4–0.7 review, <0.4 drop) and skills (≥0.6 attach with prob).
  • Dedup hash — sha256(norm company+title+city) used to detect duplicate postings.
  • Same-posting score — the Jev score 0/1/2 used in dedup: 2 merge (keep earliest posted_at, all links), 1 review.
  • AI Job Index — active AI jobs expressed as an index with base 100 on day 1.
  • Source anomaly — a volume swing greater than 50%, excluded from deltas.
  • Insufficient history — the label shown until 7 days of data exist.
  • Prep resource — a resource collected from YouTube Data API, Dev.to, HN Algolia, arXiv, GitHub search, or RSS, stored as title, URL, date, metadata, and a short summary in our own words.
  • Rerank score — the Jev score per (JD, resource) pair: not / somewhat / very useful.
  • ALERT_MIN_FIT — the fit threshold above which new jobs appear in the daily digest.
  • ingestion_runs — the log of pipeline runs.
  • api_usage — the log of service, calls, and tokens.
  • decision_cache — the cache keyed by sha256(state+question_set_version+model).
  • NOTES.md — the file where unknowns, assumptions, probe results, query plan, batch size, and Jev limits are recorded.
/ design preview
/: Read hero and top-fit rows
Sign Up: Create identity
/: Choose to enter product
Login: Submit credentials
Login: Retry with corrected credentials
/profile: Enter skills and preferences
/profile: 1. Save profile and recompute fit
/profile: Review score-composition schematic
/profile: 2. Correct and resubmit save
/jobs: 1. Apply filters and choose sort
/jobs: 2. Clear filters
/jobs/ id : 1. Review JD and fit breakdown
/jobs/ id : Open apply link
/jobs/ id : 2. Request interview questions
/jobs/ id : 3. Retry question generation
/jobs/ id : Save and track application
/applications: 1. Update status and notes
/applications: 2. Retry save after failure
/prep: 4. Open ranked prep resource
/prep: 5. Open job prep list
/growth: 1. Read top gaps and 4-week plan
/skills/ name : 2. Read trend and share breakdown
/skills/ name : 3. Open ranked resources
/market: 4. Read index chart and gainers
/market: 5. Open skill or location context
/jobs/ id : Open job from digest
/growth: Act on rising gap skill
/ design preview
/: Read hero and top-fit rows
Sign Up: Create identity
/: Choose to enter product
Login: Submit credentials
Login: Retry with corrected credentials
/profile: Enter skills and preferences
/profile: 1. Save profile and recompute fit
/profile: Review score-composition schematic
/profile: 2. Correct and resubmit save
/jobs: 1. Apply filters and choose sort
/jobs: 2. Clear filters
/jobs/ id : 1. Review JD and fit breakdown
/jobs/ id : Open apply link
/jobs/ id : 2. Request interview questions
/jobs/ id : 3. Retry question generation
/jobs/ id : Save and track application
/applications: 1. Update status and notes
/applications: 2. Retry save after failure
/prep: 4. Open ranked prep resource
/prep: 5. Open job prep list
/growth: 1. Read top gaps and 4-week plan
/skills/ name : 2. Read trend and share breakdown
/skills/ name : 3. Open ranked resources
/market: 4. Read index chart and gainers
/market: 5. Open skill or location context
/jobs/ id : Open job from digest
/growth: Act on rising gap skill