winter-career is the working project name for the AI-Based Career Prediction System, a simple, clean, fully working AI & Data Science predictive website built as a BCA college capstone project. Its single purpose is to demonstrate the complete Data Science / AI workflow end to end — Data Collection → Data Cleaning → EDA → Visualization → Preprocessing → Feature Engineering → Machine Learning → NLP → Deep Learning → AI/LLM → Prediction — and to use that workflow to predict a suitable career/job role from a user's Education, Skills and Experience.
The website is a teaching artefact, not a commercial product. It is intended for a BCA student who builds and presents it, a project guide, and an examiner. Every page demonstrates exactly one topic in the pipeline, and the pages are connected through navigation so the whole route reads as one continuous workflow. The interface must be simple to understand, clean and professional, beginner-friendly, responsive, and free of fake buttons, placeholder pages, and hardcoded prediction results. All dataset processing and all Machine Learning prediction must be real.
The project is delivered as a Streamlit application with the files app.py, requirements.txt, README.md, and data/career_dataset.csv, and runs with:
pip install -r requirements.txt
streamlit run app.py
The system is a single Streamlit application composed of fourteen connected pages. A persistent left module rail lists the eleven workflow stages as numbered entries (01 Data Collection through 11 Prediction), each with its coded colour dot, so the pipeline is always visible and the current stage is always marked. The Home / Dashboard is the public entry surface and introduces the project, its dataset size, its feature count, the Machine Learning model in use, the model accuracy, and the workflow strip.
The application loads or accepts the sample career dataset, then carries that dataset through cleaning, exploratory analysis, visualization, preprocessing, feature engineering, model training, NLP, and a neural network demonstration. The trained Machine Learning model is the model actually used by the Career Prediction form, and the Prediction Result page presents the predicted role, confidence, recommended skills, an explanation grounded in the user's own input, and the top 3 alternative roles with probabilities.
Two accepted human personas use the system: the BCA Student / Project Presenter, who walks the pipeline stage by stage and presents the report, and the Career Seeker / Prediction User, who submits the prediction form, describes skills in the NLP text input, and asks the AI Career Assistant questions. All fourteen pages are reachable without an account; the application owns no login, no user profile, and no stored personal history.
The AI / LLM page uses an LLM API through environment variables. No API key is ever hardcoded. If no API key is available, the page shows a clear message and the rest of the website remains fully functional.
Narrow exclusions. The project does not include user accounts, authentication, role-based permissions, payment, or any commercial career-placement service. It does not include stock photography, 3D renders, or decorative illustration. It does not include any capability beyond the eleven workflow stages, the prediction form and result, and the project report.
Everything the user sees is first-party Streamlit UI owned by this application. There is no provider-owned surface, no external destination, and no headless delivery: the entire product is the running Streamlit app, and every page is reachable from the navigation rail and the Home dashboard.
The current delivery horizon covers all fourteen pages and the full pipeline. The only external dependency is the optional LLM API used by the AI Career Assistant; that dependency is read from environment variables at runtime, and its absence degrades exactly one page's assistant responses while leaving every other page and every other capability working. The dataset is a local CSV file shipped with the project and can also be replaced by a user-uploaded CSV on the Data Collection page.
Nothing in this document is deferred to a future release. The Project Report page's "Future Scope" section is report content describing possible extensions, not a commitment to build them.
No reference directive with content_source authority was supplied, so no source content inventory is included. The sample career dataset is defined by the user's own field list and is specified in the Functional Requirements and Page Content coverage below.
data/career_dataset.csv is shown so the page is never blank; success — the dataset renders with its counts, column names, and preview rows; error — a malformed or unreadable CSV produces a clear message naming the parse failure and the sample dataset remains available; recovery — the user can re-upload a corrected file or continue with the sample dataset.FR-01 — Project identity and purpose (explicit) As a BCA Student / Project Presenter, I should have a simple, clean, fully working AI & Data Science predictive website titled AI-Based Career Prediction System whose single purpose is to demonstrate the complete Data Science / AI workflow and use it to predict a suitable career/job role, so that the project reads as one coherent capstone rather than a collection of unrelated demos. Lifecycle: initiated by the presenter opening the app; observable result is the Home / Dashboard masthead and workflow strip; failure is a dataset or model summary that cannot be computed, reported on the dashboard with a retry; continuation is navigation into any module. Access: none. Owner: Home / Dashboard.
FR-02 — Complete workflow demonstration (explicit) As a BCA Student / Project Presenter, I should be able to walk the pipeline Data Collection → Data Cleaning → EDA → Visualization → Preprocessing → Feature Engineering → Machine Learning → NLP → Deep Learning → AI/LLM → Prediction, with each page demonstrating exactly one topic, so that the examiner can see every stage of the workflow. Lifecycle: initiated from the module rail or the workflow strip; observable result is the selected stage's page; failure is a stage whose prerequisite data is missing, reported on that page with a link to the prerequisite; continuation is the next stage in the rail. Access: none. Owner: the eleven stage pages plus Home / Dashboard.
FR-03 — Prediction inputs (explicit) As a Career Seeker / Prediction User, I should be able to enter my Education, Skills and Experience — together with Specialization, Projects, and Certification — and receive a predicted suitable career/job role, so that the prediction reflects my own background. Lifecycle: initiated on Career Prediction; observable result is the predicted role on Prediction Result; failure is an incomplete form or an unavailable trained model, reported with the specific missing item; continuation is editing the form and predicting again. Access: none. Owner: Career Prediction → Prediction Result.
FR-04 — Home / Dashboard content (explicit) As a BCA Student / Project Presenter, I should see on the Home / Dashboard the project title, a short project description, the number of records, the number of features, the Machine Learning model used, the model accuracy, navigation to all modules, and the workflow Data → Cleaning → Analysis → ML → AI → Prediction, so that the project's scope and current model state are visible at a glance. Lifecycle: initiated by opening the app; observable result is the populated stat band and workflow strip; failure is an unavailable dataset or model summary, reported explicitly; continuation is navigation to any module. Access: none. Owner: Home / Dashboard.
FR-05 — Data Collection capabilities (explicit) As a BCA Student / Project Presenter, I should be able to upload a CSV file, view the uploaded dataset, see the number of rows and columns, see the column names, see the first 5/10 records, and download the dataset, so that the collection stage is demonstrable and reproducible. Lifecycle: initiated by uploading a CSV or accepting the sample dataset; observable result is the rendered dataset with counts, column names, and preview rows; failure is an unreadable CSV, reported with the parse reason while the sample dataset remains available; continuation is proceeding to Data Cleaning. Access: none. Owner: Data Collection.
FR-06 — Sample career dataset (explicit)
As a BCA Student / Project Presenter, I should have a sample career dataset containing Education, Specialization, Skills, Programming Level, Experience, Projects, Certifications, and Career/Job Role, so that every downstream stage has real data to process.
Lifecycle: initiated by the application loading data/career_dataset.csv; observable result is the dataset available on Data Collection and usable by every later stage; failure is a missing or unreadable file, reported on Data Collection with the option to upload a replacement; continuation is the cleaning stage. Access: none. Owner: Data Collection (with the file shipped in the project).
FR-07 — Data Cleaning capabilities (explicit) As a BCA Student / Project Presenter, I should see missing values, duplicate records, incorrect data types, missing value handling, duplicate removal, and the cleaned dataset, with before/after statistics in the form Before Cleaning — Rows, Missing values, Duplicates and After Cleaning — Rows, Missing values, Duplicates, so that the cleaning stage is visibly evidenced. Lifecycle: initiated by running cleaning; observable result is the two statistic blocks and the cleaned dataset; failure is a cleaning operation that cannot complete, reported by operation name with the raw dataset preserved; continuation is EDA on the cleaned dataset. Access: none. Owner: Data Cleaning.
FR-08 — EDA capabilities (explicit) As a BCA Student / Project Presenter, I should see dataset statistics — mean, median, minimum, maximum, standard deviation — plus the most common education, most common skill, most common career, and useful insights such as "Python is one of the most common skills among Data Analyst records.", so that the dataset is understood before modeling. Lifecycle: initiated by opening EDA after cleaning; observable result is the statistics table, most-common values, and insights; failure is a statistic that cannot be computed, reported by name with the others preserved; continuation is Data Visualization. Access: none. Owner: Exploratory Data Analysis (EDA).
FR-09 — Data Visualization capabilities (explicit) As a BCA Student / Project Presenter, I should see charts for Education distribution, Skills distribution, Career distribution, Experience distribution, Certification distribution, Education vs Career, and Skills vs Career, built with Matplotlib and Plotly, updating based on the selected dataset, so that the distributions and relationships are visible. Lifecycle: initiated by selecting a chart; observable result is the chart rendered against the currently selected dataset; failure is a chart that cannot be built, reported by chart name while the others remain selectable; continuation is Data Preprocessing. Access: none. Owner: Data Visualization.
FR-10 — Data Preprocessing capabilities (explicit) As a BCA Student / Project Presenter, I should see categorical encoding, numerical feature processing, missing value handling, feature scaling where required, and the train/test split, with brief explanations such as Education → One Hot Encoding, Experience → Numerical Encoding, and Skills → Multi-label Encoding, so that the preparation for Machine Learning is explicit. Lifecycle: initiated by running preprocessing; observable result is the encodings, scaling summary, and split row counts; failure is a failing step named explicitly with the cleaned dataset untouched; continuation is Feature Engineering. Access: none. Owner: Data Preprocessing.
FR-11 — Feature Engineering capabilities (explicit) As a BCA Student / Project Presenter, I should have features created — Number of Skills, Number of Projects, Experience Score, Certification Score, Programming Skill Score — shown in a table, with the worked example Python + SQL + ML = Skill Count 3, so that the engineered inputs to the models are inspectable. Lifecycle: initiated by running feature engineering; observable result is the new-feature table; failure is a feature that cannot be computed, reported by name with the others preserved; continuation is Machine Learning. Access: none. Owner: Feature Engineering.
FR-12 — Machine Learning training and evaluation (explicit) As a BCA Student / Project Presenter, I should be able to select Random Forest or Logistic Regression, train a real model, and see the training dataset, testing dataset, accuracy, precision, recall, F1 score, confusion matrix, and trained model information, so that the model's performance is evidenced. Lifecycle: initiated by selecting a model and training; observable result is the five metrics and the confusion matrix for the selected model; failure is a training failure named by model with any previously trained model preserved; continuation is using the trained model for prediction. Access: none. Owner: Machine Learning.
FR-13 — Trained model used for prediction (explicit) As a Career Seeker / Prediction User, I should have my Career Prediction scored by the Machine Learning model actually trained on the Machine Learning page, so that no prediction result is hardcoded. Lifecycle: initiated by pressing \xf0\x9f\x94\xae PREDICT CAREER; observable result is a prediction produced by the trained model; failure is an unavailable trained model, reported with a link to Machine Learning; continuation is training a model and predicting again. Access: none. Owner: Career Prediction, supported by Machine Learning.
FR-14 — NLP text processing (explicit) As a Career Seeker / Prediction User, I should be able to enter my skills or career interest in the text input labelled "Enter your skills or career interest:" — for example "I know Python, SQL and data visualization." — and have it processed using text cleaning, tokenization, and TF-IDF to identify relevant skills/career keywords, so that free text becomes a detected-skills list such as Python, SQL, Data Visualization. Lifecycle: initiated by submitting text; observable result is the Detected Skills list; failure is text yielding no recognizable skills, reported clearly rather than as a silent empty list; continuation is entering different text or moving to the AI Career Assistant. Access: none. Owner: NLP.
FR-15 — Sentiment analysis demonstration (explicit) As a BCA Student / Project Presenter, I should see a simple sentiment analysis demonstration using sample feedback data on the NLP page, so that the NLP stage covers both keyword extraction and sentiment. Lifecycle: initiated by opening the NLP page; observable result is the sentiment results over the sample feedback data; failure is a sentiment computation that cannot complete, reported clearly; continuation is the Deep Learning stage. Access: none. Owner: NLP.
FR-16 — Neural network demonstration (explicit) As a BCA Student / Project Presenter, I should see a simple neural network built with Scikit-learn MLPClassifier (or TensorFlow/Keras if appropriate) using the processed career dataset, showing the neural network architecture, training accuracy, testing accuracy, and a loss/accuracy graph if available, kept simple for a BCA academic project. Lifecycle: initiated by training the network; observable result is the architecture, both accuracies, and the loss/accuracy graph; failure is a training failure reported with the graph explicitly marked unavailable; continuation is the AI / LLM stage. Access: none. Owner: Deep Learning / Neural Network.
FR-17 — AI Career Assistant (explicit) As a Career Seeker / Prediction User, I should be able to ask the AI Career Assistant questions such as "Which skills should I learn for Data Science?", "What career can I choose after BCA?", and "How can I improve my Python skills?", and receive an answer from an LLM API, so that generative AI guidance is demonstrated. Lifecycle: initiated by submitting a question; observable result is the assistant's answer; failure is an API error reported clearly; continuation is asking another question. Access: none. Owner: AI / LLM.
FR-18 — LLM API key handling (explicit) As a BCA Student / Project Presenter, I should have the LLM API key read from environment variables with no key hardcoded anywhere, and when no API key is available I should see a clear message while the rest of the website remains fully functional, so that the project is safe to share and still demonstrable without a key. Lifecycle: initiated by the application checking the environment at runtime; observable result is either a working assistant or the clear no-key message; failure is a missing key, which is exactly the handled case; continuation is every other page continuing to work normally. Access: none. Owner: AI / LLM.
FR-19 — Career Prediction form (explicit) As a Career Seeker / Prediction User, I should complete a form with Education (10th, 12th, Diploma, BCA, B.Tech, MCA, Other), Specialization (Computer Science, IT, Data Science, AI/ML, Software Engineering, Other), Skills as a multiple selection (Python, Java, C/C++, JavaScript, HTML/CSS, SQL, Machine Learning, Data Analysis, Data Visualization, AI, Communication, Problem Solving), Experience (Fresher, <1 Year, 1–2 Years, 2–5 Years, 5+ Years), Projects (0, 1, 2–3, 4+), and Certification (Yes, No), and then press the large \xf0\x9f\x94\xae PREDICT CAREER button, so that my background is captured exactly as the model expects. Lifecycle: initiated by completing the form; observable result is the submitted prediction request; failure is an incomplete form, reported with the missing field named; continuation is the Prediction Result. Access: none. Owner: Career Prediction.
FR-20 — Prediction Result content (explicit) As a Career Seeker / Prediction User, I should see the predicted career, the confidence percentage, recommended skills, a "Why this prediction?" explanation based on my input — for example "Your prediction is influenced by your Python, SQL and Data Analysis skills and your BCA background." — and the top 3 possible career roles with their prediction probabilities, presented without making the interface complicated, so that the result is both actionable and explainable. Lifecycle: initiated by a successful prediction; observable result is the predicted career, confidence, recommended skills, explanation, and top 3 alternatives; failure is a scoring failure reported with the form input preserved; continuation is returning to Career Prediction to predict again. Access: none. Owner: Prediction Result.
FR-21 — Project Report content and download (explicit) As a BCA Student / Project Presenter, I should have a Project Report page containing Introduction, Problem Statement, Objectives, Dataset, Data Collection, Data Cleaning, EDA, Visualization, Preprocessing, Feature Engineering, Machine Learning, NLP, Deep Learning, AI/LLM, Prediction, Results, Limitations, Future Scope, and Conclusion, with a Download Report button, so that the written submission can be read in the app and downloaded. Lifecycle: initiated by opening the report; observable result is all nineteen sections rendered and a working download; failure is a download failure reported clearly; continuation is retrying the download or navigating away. Access: none. Owner: Project Report.
FR-22 — Navigation connectivity (explicit) As a BCA Student / Project Presenter, I should be able to reach every page through navigation, with the eleven workflow stages listed as numbered entries in the module rail and the current stage marked, so that no page is orphaned and the pipeline reads as one route. Lifecycle: initiated from the rail or the workflow strip; observable result is the destination page with the current stage marked; failure is a stage whose prerequisite is missing, reported on that page with a link to the prerequisite; continuation is the next stage. Access: none. Owner: Home / Dashboard and the shared module rail.
FR-23 — Dataset availability before analysis (required_inference)
As a BCA Student / Project Presenter, I should have the sample career dataset loaded or provided before any analysis or modeling stage runs, so that the accepted pipeline is executable from the first stage.
Lifecycle: initiated by the application loading data/career_dataset.csv or by a CSV upload; observable result is a dataset available to every downstream stage; failure is a missing or unreadable dataset, reported on Data Collection with the upload alternative; continuation is Data Cleaning. Access: none. Owner: Data Collection.
FR-24 — Ordered processing before prediction (required_inference) As a BCA Student / Project Presenter, I should have the selected dataset processed through cleaning, preprocessing, feature engineering, and model training before the final career prediction is available, so that the prediction is genuinely produced by the demonstrated workflow. Lifecycle: initiated by running each stage in order; observable result is a trained model ready for scoring; failure is a stage that has not been run, reported on Career Prediction with a link to the missing stage; continuation is completing that stage and predicting. Access: none. Owner: the pipeline stages, surfaced on Career Prediction.
FR-25 — Trained model reuse (required_inference) As a Career Seeker / Prediction User, I should have the trained machine-learning model reused for my Career Prediction submission and Prediction Result, so that the result reflects the model I can inspect on the Machine Learning page. Lifecycle: initiated by pressing \xf0\x9f\x94\xae PREDICT CAREER; observable result is a prediction attributed to the currently trained model; failure is an unavailable model, reported with a link to Machine Learning; continuation is training and predicting again. Access: none. Owner: Career Prediction → Prediction Result.
FR-26 — Optional LLM key with functional fallback (required_inference) As a BCA Student / Project Presenter, I should have the LLM API key treated as optional and read only from environment variables, with the AI / LLM page remaining functional by clearly reporting unavailable assistance when no key is present, so that the project runs end to end without any secret. Lifecycle: initiated by the runtime environment check; observable result is either assistant answers or the clear unavailable message; failure is the missing key, which is handled rather than fatal; continuation is every other page working normally. Access: none. Owner: AI / LLM.
FR-27 — Runnable project files and dependencies (required_inference)
As a BCA Student / Project Presenter, I should have app.py, requirements.txt, README.md, and data/career_dataset.csv present so that the project runs with pip install -r requirements.txt and streamlit run app.py, so that the examiner can reproduce the demonstration.
Lifecycle: initiated by installing dependencies and starting the app; observable result is the running application on Home / Dashboard; failure is a missing dependency or file, reported by the install or start command; continuation is the full pipeline walkthrough. Access: none. Owner: the project itself (no page).
Product context. This persona is building and demonstrating the AI-Based Career Prediction System as a BCA college project. They are the person who runs the app in front of a guide or examiner, and they need the whole Data Science / AI pipeline to be visible, ordered, and honest — a working instrument rather than a pitch.
Primary goal. Walk the complete workflow — Data Collection, Data Cleaning, EDA, Visualization, Preprocessing, Feature Engineering, Machine Learning, NLP, Deep Learning, AI/LLM, Prediction — and present the Project Report, so that every stage of the capstone is demonstrably real.
Distinct accepted responsibilities. Uploading or accepting the sample career dataset and inspecting its shape, columns, and preview rows; running cleaning and reading the before/after statistics; reading the EDA statistics and most-common values; selecting and reading each visualization; running preprocessing and reading the encoding explanations; running feature engineering and reading the new-feature table; selecting Random Forest or Logistic Regression and reading accuracy, precision, recall, F1 score, and the confusion matrix; entering text on the NLP page and reading the detected skills and sentiment demonstration; training the neural network and reading its architecture, accuracies, and loss/accuracy graph; reading the AI Career Assistant's availability status; and reading and downloading the Project Report.
Relevant inputs or decisions. Which CSV to upload or whether to use the sample; which chart to view; which model to select; which text to submit on the NLP page; whether to train the neural network; whether to download the report.
Interactions with other accepted participants. The presenter prepares the dataset and the trained model that the Career Seeker / Prediction User depends on, and can demonstrate the prediction form and result on the seeker's behalf during a presentation.
Observable success. Every stage renders real output from the real dataset, the metrics and charts are populated, the report downloads, and no page is a placeholder.
Product context. This persona wants a concrete answer to "which career fits my background?" and arrives at the Career Prediction page with their own education, specialization, skills, experience, projects, and certification in mind. They may also describe their skills in free text and ask the assistant for guidance.
Primary goal. Submit their Education, Specialization, Skills, Experience, Projects, and Certification, press \xf0\x9f\x94\xae PREDICT CAREER, and receive a predicted career role with confidence, recommended skills, an explanation they can understand, and the top 3 alternatives.
Distinct accepted responsibilities. Completing the six-field prediction form with the exact allowed values; pressing \xf0\x9f\x94\xae PREDICT CAREER; reading the predicted career and confidence; reading the recommended skills; reading the "Why this prediction?" explanation; reading the top 3 possible career roles with probabilities; entering skills or career interest text on the NLP page and reading the detected skills; and asking the AI Career Assistant career questions.
Relevant inputs or decisions. Their own education level, specialization, multi-selected skills, experience band, project count, and certification status; the free-text description of their skills; the questions they ask the assistant.
Interactions with other accepted participants. The seeker depends on the BCA Student / Project Presenter having loaded the dataset and trained the model; if no model is trained, the seeker is directed to the Machine Learning page rather than shown a fabricated result.
Observable success. A predicted career appears with a confidence percentage, recommended skills, an explanation that names their own inputs, and three ranked alternatives — all produced by the trained model, never hardcoded.
pip install -r requirements.txt and streamlit run app.py, then opens the app in a browser.data/career_dataset.csv and computes the dataset summary and the current model summary.data/career_dataset.csv is already displayed, so the page is never blank.The visual direction is typographic infrastructure for a career prediction engine — wayfinding clarity, warm ground, signal-colour pipeline — after Erik Spiekermann. The project is a teaching artefact, so it reads like a well-set technical manual: numbered, gridded, honest, and warm enough to sit with for an hour. Typography is treated as infrastructure, and the eleven pipeline stages are treated like transit lines.
Colour tokens (light mode).
| Role | Hex | Use |
|---|---|---|
| Background | #F4EFE6 | Warm paper ground; never pure white |
| Surface | #FFFBF3 | Even warmer card surface |
| Text | #1C1A17 | Warm near-black ink, ≈14:1 contrast at body size |
| Primary | #C2410C | Burnt signal orange: active pipeline stage, PREDICT CAREER button, selected model |
| Accent | #1F6F5C | Deep transit green: confirmed/positive states — accuracy figures, "cleaned", "certified" |
| Muted | #8A8175 | Labels, metadata, rules |
| Border rule | #E2D9C9 | 1px panel and table borders |
| Coded wayfinding — red | #C2410C | 4px rules, node dots, small caps labels only |
| Coded wayfinding — yellow | #D9A400 | 4px rules, node dots, small caps labels only |
| Coded wayfinding — green | #1F6F5C | 4px rules, node dots, small caps labels only |
| Coded wayfinding — blue | #2A5C8A | 4px rules, node dots, small caps labels only; never the primary or button colour |
Each of the eleven workflow modules carries one of the four coded tag colours, used only as 4px rules, node dots, and small caps labels — never as large fills. Blue appears only as one of the four coded wayfinding accents.
Typography. Headings use Fira Sans at 700 and 800 for display headings with tight tracking (−0.02em) and sentence case for page titles, and 500 for section headings. A second voice, Fira Mono, carries all data, metrics, column names, code, and pipeline labels, set in uppercase with +0.08em tracking so numbers read like a timetable. Body text is Fira Sans. Headings are flush-left and ragged-right, never centred.
Type scale (1.250 modular ratio): 64 / 48 / 34 / 26 / 20 / 17 / 15 / 13. Display page titles use clamp(40px, 7vw, 64px) from mobile to desktop; section headings 26–34px; body 17px with 1.6 leading; mono labels 13px uppercase. Measure is capped at 68ch for report prose and 46ch for captions.
Shape language. Rectilinear and honest. Maximum 2px corner radius on cards and 6px on buttons — no pill shapes, no blobs, no soft-shadow float. Every panel is a bordered rectangle with a 1px rule in #E2D9C9 and a 3px top edge in its stage's coded colour. Rules are drawn, not implied: horizontal hairlines separate every label/value pair, every table row, and every metric block. Pipeline node dots are 10px flat circles with no glow.
Layout. A visible 12-column grid with a persistent left rail — 240px on desktop, collapsing to a horizontal scrollable module strip on mobile — listing the eleven workflow stages as numbered entries, 01 Data Collection through 11 Prediction, each with its coded colour dot and the current stage marked by a burnt-orange 3px left rule and a filled dot. Content sits in a single wide column with a right-hand metadata gutter on desktop for row counts, feature counts, and model stats. The Home page leads with a full-width workflow strip: eleven numbered nodes connected by a 2px hairline, horizontally scrollable on mobile with each node fully readable as it passes. Tables are the primary visual object — ruled, tabular numerals, alternating warm-paper row tints.
Imagery. Diagrammatic and documentary, never decorative. The Matplotlib and Plotly charts are the imagery, styled to the palette: warm paper plot backgrounds, ink axes, coded series colours, no default chart chrome. Supporting visuals are schematic — the pipeline diagram, an encoding table, a confusion matrix rendered as a ruled grid, a TF-IDF token list set in Fira Mono. No stock photography, no 3D renders, no illustration for its own sake; the interface itself is the visual content.
Avoid. Any blue or indigo as the primary or button colour — #0057FF, #2563EB, #4F46E5, #6366F1, #7C3AED and neighbours are banned; blue appears only as one of four small coded wayfinding accents. Pure white grounds and pure black text. Inter, Roboto, Arial, Helvetica, Open Sans, Lato, Poppins, and system-ui for headings or body — Fira Sans and Fira Mono only, with monospace reserved for data and labels. Gradient-blob heroes, glassmorphism, frosted panels, and decorative blur. Grids of identical hover-lift cards with soft drop shadows — panels are bordered, ruled, and flat. Pill-shaped buttons, blob shapes, and corner radii above 6px. Centred hero stacks and decorative looping animation or parallax. Stock photography, 3D renders, and illustration used as decoration rather than as a diagram. The generic indigo/blue-on-white SaaS template is forbidden.
The Home hero is a full-width, left-aligned typographic masthead on warm paper (#F4EFE6) — not a centred SaaS stack. A 13px Fira Mono uppercase kicker reads BCA CAPSTONE · AI & DATA SCIENCE above a 64px Fira Sans 800 headline, AI-Based Career Prediction System, that spans the viewport width and wraps to two lines on desktop and three on mobile at clamp(40px, 7vw, 64px). Directly beneath, a single 46ch paragraph of body copy explains the one purpose of the project. Below that sits a horizontal ruled stat band of four label/value pairs in Fira Mono — RECORDS 500 · FEATURES 8 · MODEL RANDOM FOREST · ACCURACY 87.4% — separated by 1px vertical rules, with the numbers counting up once on first paint.
The signature element is the eleven-node workflow strip running edge to edge beneath the stat band: each node is a numbered Fira Mono label with its coded colour dot, connected by a 2px hairline, with the final PREDICTION node set in burnt orange and brighter than the rest. On load, the active stage's connector draws left-to-right over 400ms. On mobile the strip is horizontally scrollable (overflow-x: auto) so every node becomes fully readable as it passes. There is no gradient, no blob, and no blue button; the only filled element on the screen is the burnt-orange PREDICT CAREER control in the top-right of the content column.
The concept recomposes only accepted content and controls — the project title, the description, the four dashboard statistics, the workflow stages, and the navigation into them. It introduces no new behaviour, page, or destination.
Interaction Model: Static (direction) Motion Tempo: restrained Hero Dimensionality: flat
Motion is functional and purposeful, in the spirit of transit signage: 160–220ms ease-out on state changes, no bounce, no parallax, no decorative loops. The one expressive moment is the workflow strip, where the active stage's node fills with its coded colour and its connector rule draws left-to-right over 400ms when a page loads. Metric numbers count up once on first paint (600ms, ease-out) and then hold still. Hover on a module rail entry moves the burnt-orange rule 4px; nothing lifts or scales.
Landing Hero Motion Brief
prefers-reduced-motion, the connector rule is drawn immediately at full length, metric values appear at their final figures without counting up, and the workflow strip wraps into rows or sits in a horizontally scrollable row (overflow-x: auto) whose further nodes are reached by scrolling — every node fully readable, nothing cut off.No user-requested 3D or WebGL hero was specified, and the direction's hero dimensionality is flat, so no Canvas, R3F, or Drei scene is required.
NFR-01 — Fully working (explicit) Every page must render real output from the real dataset. No page may be a placeholder, and no control may be a fake button. Rationale: the source states the website must be fully working with no fake buttons and no placeholder pages.
NFR-02 — No hardcoded prediction results (explicit) Prediction results must be produced by the trained Machine Learning model from the user's submitted inputs. No predicted career, confidence value, or probability may be hardcoded. Rationale: explicit source constraint.
NFR-03 — Real dataset processing and real ML prediction (explicit) The dataset must be genuinely parsed, cleaned, analyzed, visualized, preprocessed, and used to train models; the prediction must come from that trained model. Rationale: explicit source constraint.
NFR-04 — No hardcoded API key (explicit) The LLM API key must be read from environment variables only. No key may appear in source, configuration committed to the repository, or the README. Rationale: explicit source constraint.
NFR-05 — Functional without an API key (explicit) When no API key is available, the AI / LLM page must show a clear message and the rest of the website must remain fully functional. Rationale: explicit source constraint.
NFR-06 — Simple and beginner-friendly (explicit) The project must stay simple and understandable for a BCA student: readable code, plain labels, and no unnecessary abstraction. Rationale: explicit source constraint.
NFR-07 — Clean and professional (explicit) The interface must be clean and professional, following the typographic and colour direction in sections 6–8. Rationale: explicit source constraint.
NFR-08 — Responsive (explicit) The website must be responsive. At 375px, 768px, and 1280px, headlines, wordmarks, labels, numbers, cards, and controls must stay entirely inside the viewport and their container, wrapping or scaling to fit, and no other element may cover any part of them. The left module rail collapses to a horizontal scrollable module strip on mobile. The workflow strip is horizontally scrollable on mobile with each node fully readable as it passes. Rationale: explicit source constraint plus the direction's readability rule.
NFR-09 — Navigation connectivity (explicit) All pages must be connected through navigation, with the eleven workflow stages listed as numbered entries and the current stage marked. Rationale: explicit source constraint.
NFR-10 — Reproducible run (explicit)
The project must run with pip install -r requirements.txt followed by streamlit run app.py, using the files app.py, requirements.txt, README.md, and data/career_dataset.csv. Rationale: explicit source constraint.
NFR-11 — Single clear purpose (explicit) The final project must have one clear purpose — use the complete Data Science/AI workflow to build a Career Prediction System — while each page demonstrates one specific topic. Rationale: explicit source constraint.
NFR-12 — Accessible contrast and legibility (required_inference)
Body text at #1C1A17 on #F4EFE6 must retain its high contrast (≈14:1), and mono labels must remain legible at 13px uppercase with +0.08em tracking. Rationale: required to make the direction's stated contrast and label sizing actually hold in the delivered interface.
All choices below are explicit user requirements.
app.py.MLPClassifier, or TensorFlow/Keras if appropriate.app.py, requirements.txt, README.md, data/career_dataset.csv.pip install -r requirements.txt then streamlit run app.py.No container, orchestration, or deployment tooling is required by the source; the project runs locally as a Streamlit application.
Constraints (binding).
pip install -r requirements.txt and streamlit run app.py.Assumptions.
data/career_dataset.csv is shipped with the project and contains the eight columns Education, Specialization, Skills, Programming Level, Experience, Projects, Certifications, and Career/Job Role. The source specifies the columns but not the row count; the illustrative figures in the source (500 rows before cleaning, 492 after) are examples of the before/after display, not a mandated dataset size.Presentation and technology defaults.
data/career_dataset.csv with the columns Education, Specialization, Skills, Programming Level, Experience, Projects, Certifications, and Career/Job Role.No completed page designs yet.
Completed design pages will appear here when they are ready to preview.
No completed page designs yet.
Completed design pages will appear here when they are ready to preview.
No comments yet. Be the first!