AI-Based Career Prediction System is a simple, clean, fully working AI & Data Science predictive website built as a BCA college project. Its single purpose is to demonstrate the complete Data Science / AI workflow end to end:
Data Collection → Data Cleaning → EDA → Visualization → Preprocessing → Feature Engineering → Machine Learning → NLP → Deep Learning → AI/LLM → Prediction
Each page of the website demonstrates one specific topic in that pipeline, and the pipeline culminates in a real prediction: the user's Education, Skills and Experience are used to predict a suitable career/job role.
The audience is the BCA student who authors and demonstrates the project, the project guide, and the college examiner. The website must read as real data science rather than a toy: real dataset processing, real Machine Learning prediction, no hardcoded prediction results, no fake buttons, and no placeholder pages. It must remain simple to understand, clean and professional, beginner-friendly, and responsive.
The system is a single Python/Streamlit application (app.py) that runs locally with:
pip install -r requirements.txt
streamlit run app.py
It ships with four files: app.py, requirements.txt, README.md, and data/career_dataset.csv. The bundled sample career dataset contains the columns Education, Specialization, Skills, Programming Level, Experience, Projects, Certifications, Career/Job Role, and it is the default dataset for every workflow module. A user may also upload their own CSV on the Data Collection page, and every downstream module then operates on that selected dataset.
All fourteen pages are connected through navigation. The left rail lists the eleven numbered pipeline modules (01–11) plus the Home / Dashboard, Career Prediction, Prediction Result, and Project Report destinations, so the student always sees where they are in the pipeline.
Actors:
Non-persona actors: the LLM API provider (external, accessed only through environment variables) and the application runtime (Streamlit/Python process performing dataset processing, model training and prediction).
Narrow exclusions: no hardcoded API key; no hardcoded prediction results; no fake buttons; no placeholder pages; no blue/indigo SaaS template look; no stock photography, 3D, gradients, glassmorphism or dark mode.
Everything in this project is delivered as first-party custom Streamlit UI owned by the application. There is no application account system, no login, no signup, and no per-user stored profile: the source never asks for one, and every accepted journey — uploading a dataset, walking the pipeline, filling the prediction form, reading the result, asking the assistant — is a single-session, anonymous interaction. The Home / Dashboard is the anonymous product entry and every page is reachable without identity.
The only external dependency is the LLM API used by the AI Career Assistant. It is provider-owned and reached exclusively through an environment variable. If no API key is available, the AI / LLM page shows a clear message and the rest of the website remains fully functional. No API key is ever hardcoded.
Current scope is the full eleven-stage pipeline plus the prediction and report destinations. There is no future-horizon section in the source; nothing beyond the accepted pages is committed.
Not applicable — no reference directive declares content_source.
FR-1 — Pipeline demonstration (explicit). As a Project Demonstrator / Evaluator, I should be able to walk the complete workflow Data Collection → Data Cleaning → EDA → Visualization → Preprocessing → Feature Engineering → Machine Learning → NLP → Deep Learning → AI/LLM → Prediction across the website, so that each page demonstrates one specific Data Science/AI topic. Lifecycle: initiator = demonstrator; trigger = opening the app; observable result = each stage page renders its own real output; failure/recovery = a stage that cannot run names the missing prerequisite and links back to it; continuation = the next stage in the pipeline. Acceptance: all eleven stages are reachable and each produces real output from the selected dataset.
FR-2 — Home / Dashboard (explicit). As a Career Seeker / Student User or Project Demonstrator / Evaluator, I should see the project title, a short project description, the number of records, the number of features, the Machine Learning model used, the model accuracy, navigation to all modules, and the workflow Data → Cleaning → Analysis → ML → AI → Prediction. Lifecycle: initiator = visitor; trigger = opening the app; observable result = dashboard statistics and workflow strip; failure/recovery = unavailable statistics are labelled rather than invented; continuation = navigate to any module. Acceptance: every listed item is visible and every module is reachable.
FR-3 — Data Collection (explicit). As a Project Demonstrator / Evaluator, I should upload a CSV file, view the uploaded dataset, see the number of rows and columns, see the column names, see the first 5/10 records, and download the dataset. Lifecycle: initiator = demonstrator; trigger = CSV upload or default sample load; observable result = dataset preview with row/column counts, column names and leading records; failure/recovery = malformed CSV shows a clear message and keeps the previous dataset; continuation = proceed to Data Cleaning. Acceptance: the bundled sample career dataset contains Education, Specialization, Skills, Programming Level, Experience, Projects, Certifications, Career/Job Role, and the download produces the selected dataset.
FR-4 — Data Cleaning (explicit). As a Project Demonstrator / Evaluator, I should see missing values, duplicate records and incorrect data types, apply missing value handling and duplicate removal, view the cleaned dataset, and see before/after statistics. Lifecycle: initiator = demonstrator; trigger = run cleaning; observable result = before/after rows, missing values and duplicates (e.g. Before: Rows 500, Missing values 25, Duplicates 8 → After: Rows 492, Missing values 0, Duplicates 0) plus the cleaned dataset; failure/recovery = an incompatible column is named and the raw dataset is preserved; continuation = proceed to EDA. Acceptance: before/after statistics and the cleaned dataset are both displayed.
FR-5 — Exploratory Data Analysis (explicit). As a Project Demonstrator / Evaluator, I should see dataset statistics (mean, median, minimum, maximum, standard deviation), the most common education, the most common skill, the most common career, and useful dataset insights. Lifecycle: initiator = demonstrator; trigger = open EDA with a selected dataset; observable result = statistics, most-common values and insight statements such as "Python is one of the most common skills among Data Analyst records"; failure/recovery = non-numeric columns are excluded from numeric statistics with a note; continuation = proceed to Visualization. Acceptance: all listed statistics and the three most-common values are shown.
FR-6 — Data Visualization (explicit). As a Project Demonstrator / Evaluator, I should see charts for Education distribution, Skills distribution, Career distribution, Experience distribution, Certification distribution, Education vs Career, and Skills vs Career, built with Matplotlib and Plotly, updating based on the selected dataset. Lifecycle: initiator = demonstrator; trigger = open Visualization or change the selected dataset; observable result = the seven chart types re-render for the current dataset; failure/recovery = a chart that cannot be built is reported while the others still render; continuation = proceed to Preprocessing. Acceptance: all seven chart types exist and update when the selected dataset changes.
FR-7 — Data Preprocessing (explicit). As a Project Demonstrator / Evaluator, I should see categorical encoding, numerical feature processing, missing value handling, feature scaling where required, and the train/test split, with brief explanations of what preprocessing is doing (e.g. Education → One Hot Encoding, Experience → Numerical Encoding, Skills → Multi-label Encoding). Lifecycle: initiator = demonstrator; trigger = run preprocessing; observable result = encoding summary, scaling note and train/test split sizes; failure/recovery = an unencodable column is named; continuation = proceed to Feature Engineering. Acceptance: each preprocessing step is shown with its explanation.
FR-8 — Feature Engineering (explicit). As a Project Demonstrator / Evaluator, I should create Number of Skills, Number of Projects, Experience Score, Certification Score and Programming Skill Score, and see the new features in a table. Lifecycle: initiator = demonstrator; trigger = run feature engineering; observable result = engineered-feature table (e.g. Python + SQL + ML = Skill Count 3); failure/recovery = a missing source column is named explicitly; continuation = proceed to Machine Learning. Acceptance: all five engineered features appear in the table.
FR-9 — Machine Learning (explicit). As a Project Demonstrator / Evaluator, I should select Random Forest or Logistic Regression, train a real model, and see the training dataset, testing dataset, accuracy, precision, recall, F1 score, confusion matrix and trained model information. Lifecycle: initiator = demonstrator; trigger = model selection and training; observable result = metrics and confusion matrix for the selected model; failure/recovery = training failure shows a clear message and the model can be switched; continuation = the trained model is used for the final career prediction. Acceptance: both models are selectable, all six evaluation outputs are shown, and the trained model actually drives prediction.
FR-10 — NLP (explicit). As a Career Seeker / Student User, I should type into the input "Enter your skills or career interest:" (e.g. "I know Python, SQL and data visualization."), have the text processed with text cleaning, tokenization and TF-IDF, and see the relevant skills/career keywords under Detected Skills (e.g. Python, SQL, Data Visualization). Lifecycle: initiator = student; trigger = submitting free text; observable result = detected skills list; failure/recovery = empty or unprocessable input shows a clear prompt; continuation = use the detected skills to inform the Career Prediction form. Acceptance: the three processing steps are applied and the detected skills are displayed.
FR-11 — NLP sentiment demonstration (explicit). As a Project Demonstrator / Evaluator, I should see a simple sentiment analysis demonstration using sample feedback data. Lifecycle: initiator = demonstrator; trigger = open the sentiment demonstration; observable result = sentiment result over the sample feedback data; failure/recovery = unavailable sample feedback shows a clear message; continuation = return to the NLP keyword flow. Acceptance: the sentiment demonstration renders on the NLP page.
FR-12 — Deep Learning / Neural Network (explicit). As a Project Demonstrator / Evaluator, I should train a simple neural network using Scikit-learn MLPClassifier (or TensorFlow/Keras if appropriate) on the processed career dataset and see the neural network architecture, training accuracy, testing accuracy, and a loss/accuracy graph if available. Lifecycle: initiator = demonstrator; trigger = train the network; observable result = architecture, both accuracies and the graph; failure/recovery = training failure shows a clear message; continuation = compare with the Machine Learning page results. Acceptance: the section stays simple for a BCA academic project and shows all four listed outputs.
FR-13 — AI Career Assistant (explicit). As a Career Seeker / Student User, I should ask questions such as "Which skills should I learn for Data Science?", "What career can I choose after BCA?" and "How can I improve my Python skills?" and receive answers from an LLM API accessed through environment variables. Lifecycle: initiator = student; trigger = submitting a question; observable result = generated answer; failure/recovery = if no API key is available, a clear message is shown and the rest of the website remains fully functional; continuation = ask another question or continue to Career Prediction. Acceptance: no API key is hardcoded anywhere, and the unavailable-key state is explicit.
FR-14 — Career Prediction form (explicit). As a Career Seeker / Student User, I should fill a simple form with Education (10th, 12th, Diploma, BCA, B.Tech, MCA, Other), Specialization (Computer Science, IT, Data Science, AI/ML, Software Engineering, Other), multi-select Skills (Python, Java, C/C++, JavaScript, HTML/CSS, SQL, Machine Learning, Data Analysis, Data Visualization, AI, Communication, Problem Solving), Experience (Fresher, <1 Year, 1–2 Years, 2–5 Years, 5+ Years), Projects (0, 1, 2–3, 4+) and Certification (Yes, No), then press the large \xf0\x9f\x94\xae PREDICT CAREER button. Lifecycle: initiator = student; trigger = completing the form and pressing the button; observable result = a prediction request is executed against the real trained model; failure/recovery = incomplete required input shows a clear validation message; continuation = the Prediction Result is shown. Acceptance: every listed option is present exactly as specified and the button is large and functional.
FR-15 — Prediction Result (explicit). As a Career Seeker / Student User, I should see the predicted career (e.g. Data Analyst), the confidence (e.g. 85%), recommended skills (e.g. Python, SQL, Data Analysis, Power BI), a "Why this prediction?" explanation based on my input (e.g. "Your prediction is influenced by your Python, SQL and Data Analysis skills and your BCA background"), and the top 3 possible career roles with their prediction probabilities, without the interface becoming complicated. Lifecycle: initiator = student; trigger = a completed prediction; observable result = the full result panel; failure/recovery = prediction failure shows a clear message and the form can be resubmitted; continuation = adjust inputs and predict again. Acceptance: all five result elements are shown and no result is hardcoded.
FR-16 — Project Report (explicit). As a Project Demonstrator / Evaluator, I should read a report containing Introduction, Problem Statement, Objectives, Dataset, Data Collection, Data Cleaning, EDA, Visualization, Preprocessing, Feature Engineering, Machine Learning, NLP, Deep Learning, AI/LLM, Prediction, Results, Limitations, Future Scope and Conclusion, and press a Download Report button. Lifecycle: initiator = demonstrator; trigger = opening the report or pressing download; observable result = the report content and a downloaded report file; failure/recovery = download failure shows a clear message; continuation = return to any module. Acceptance: every listed section is present and the download button produces a real file.
FR-17 — Sample dataset availability (required_inference). As a Project Demonstrator / Evaluator, I should have the sample career dataset available at data/career_dataset.csv for the workflow modules and real model prediction. Lifecycle: initiator = application runtime on first load; trigger = app start; observable result = the sample dataset is the default selection for every module; failure/recovery = a missing file shows a clear message and prompts upload on Data Collection; continuation = the pipeline runs on the uploaded dataset. Acceptance: the bundled dataset contains the eight specified columns.
FR-18 — Dataset processing prerequisite (required_inference). As a Project Demonstrator / Evaluator, I should have a submitted or sample dataset processed before dependent statistics, charts, preprocessing, feature engineering, model evaluation, neural-network and prediction views can operate. Lifecycle: initiator = application runtime; trigger = opening a dependent module; observable result = dependent modules render from the currently selected dataset; failure/recovery = a module without a usable dataset shows a clear prerequisite message and a route to Data Collection; continuation = the module renders once a dataset is selected. Acceptance: no dependent module silently fabricates data.
FR-19 — LLM key prerequisite (required_inference). As a Career Seeker / Student User, I should be able to use the AI Career Assistant only when an LLM API key is supplied through an environment variable; without it the assistant shows its unavailable-key message while the other modules remain functional. Lifecycle: initiator = student; trigger = asking a question; observable result = answer or unavailable-key message; failure/recovery = the message names the missing environment variable; continuation = the rest of the website stays usable. Acceptance: the key is read from the environment only.
FR-20 — Navigation connectivity (explicit). As a Career Seeker / Student User or Project Demonstrator / Evaluator, I should reach every page through navigation, with no fake buttons and no placeholder pages. Lifecycle: initiator = any visitor; trigger = using the module rail or in-page links; observable result = the target page opens with real content; failure/recovery = a link to an unavailable state shows the reason; continuation = the user returns to the pipeline. Acceptance: all fourteen pages are connected and every control performs a real action.
Provenance: required_inference (Planning Scope).
Product context. A BCA student or job seeker who wants to know which career role fits their Education, Specialization, Skills, Experience, Projects and Certification. They are the reason the prediction exists; the pipeline pages are how they come to trust the answer.
Primary goal. Enter their own profile once and receive a suitable career/job role with a confidence figure, recommended skills, a plain-language explanation of why, and the top 3 alternative roles.
Distinct accepted responsibilities. Completing the Career Prediction form with the exact Education, Specialization, multi-select Skills, Experience, Projects and Certification options; pressing \xf0\x9f\x94\xae PREDICT CAREER; reading the predicted career, confidence, recommended skills, "Why this prediction?" explanation and top-3 alternatives; describing their skills in free text on the NLP page ("I know Python, SQL and data visualization.") and reading the Detected Skills; asking the AI Career Assistant questions such as "Which skills should I learn for Data Science?", "What career can I choose after BCA?" and "How can I improve my Python skills?".
Relevant inputs or decisions. Which education level and specialization describe them; which of the twelve skills they actually have; their experience band; how many projects they have completed; whether they hold a certification; whether to accept the predicted role or adjust their inputs and predict again.
Interactions with other accepted participants. They consume the model that the Project Demonstrator / Evaluator trained on the Machine Learning page; the AI Career Assistant answers come from the external LLM API provider when a key is configured.
Observable success. A prediction result panel showing a predicted career, a confidence percentage, recommended skills, an explanation referencing their own inputs, and three ranked alternative roles — produced by the real trained model, never hardcoded.
Provenance: required_inference (Planning Scope).
Product context. The student author presenting the capstone to a project guide or examiner, and the evaluator assessing whether the whole Data Science/AI workflow genuinely runs. Their credibility depends on every stage showing real output from real data.
Primary goal. Demonstrate a fully working, connected pipeline in which each page proves one Data Science/AI topic on the actual dataset, and hand over a documented report.
Distinct accepted responsibilities. Uploading the sample career dataset and inspecting rows, columns, column names and the first 5/10 records; downloading the dataset; running cleaning and showing before/after statistics; reading EDA statistics and insights; rendering the seven visualization chart types; running preprocessing and explaining the encodings; creating and showing the engineered features; selecting Random Forest or Logistic Regression and presenting accuracy, precision, recall, F1 and the confusion matrix; running the NLP keyword and sentiment demonstrations; training the MLP neural network and showing architecture, training accuracy, testing accuracy and the loss/accuracy graph; demonstrating the AI Career Assistant and its unavailable-key behaviour; and presenting the Project Report with its Download Report button.
Relevant inputs or decisions. Which dataset to demonstrate with; which model to select; which chart or statistic to highlight; whether the LLM key is configured for the live demonstration.
Interactions with other accepted participants. They produce the trained model and the report that the Career Seeker / Student User depends on; they present the AI / LLM page whose answers come from the external LLM API provider.
Observable success. Every module renders real output with no placeholder pages and no fake buttons, the pipeline is navigable end to end, and the report downloads as a real file.
The creative direction is authoritative: typographic infrastructure for a career pipeline, after the muse Erik Spiekermann — typography as infrastructure, transit signage, numbered modules, data read like a departure board. Headline: an eleven-stage pipeline that reads like a public information system, warm enough for a classroom.
Colour tokens (light mode only — no dark mode).
| Role | Hex | Use |
|---|---|---|
| Background | #F2EEE5 | Warm paper ground for the whole page |
| Surface | #FFFDF7 | Cards and data panels, with a 1px #D8D1C2 rule, never a shadow |
| Text / ink | #1A1A18 | All reading text and numerals |
| Primary | #C7401F | Transit red: active module in the pipeline rail, PREDICT CAREER button, predicted-career headline, accuracy figures |
| Accent | #E8B22B | Mustard: secondary line colour for the workflow strip, chart series and stage numbers |
| Reserved line | #1F6E63 | Deep teal, used only for the neural network and LLM stages so the pipeline reads like separate transit lines |
| Muted | #6F6A5E | Captions, metadata and axis labels |
| Rule | #D8D1C2 | 1px structural rules, table row separators |
Charts use the four signal colours at full saturation on the paper ground. No gradients anywhere. Blue and indigo are forbidden (#0057FF, #2563EB, #4F46E5, #6366F1 and neighbours), as is the blue-on-white default Streamlit theme.
Typography. Headings and body: Fira Sans (Spiekermann's own humanist sans). Headings at 700–800 weight, tight -0.02em tracking, sentence case for page titles, all-caps with +0.14em tracking for module numbers, stage labels and table headers. Numerals are the ornament: model accuracy, record counts and confidence percentages are set large in 700 weight, tabular, so they read like departure-board figures. No italics, no serifs, no display face. Forbidden for headings or body: Inter, Roboto, Arial, Helvetica, Open Sans, Lato, Poppins, system-ui.
Type scale (1.25 modular). Mobile 30/24/19/16; desktop 44/32/22/17. Hero title clamp(32px, 6vw, 72px); page titles clamp(30px, 4.5vw, 44px); section titles 22px; body 17px/1.55; labels 13px all-caps; numerals clamp(36px, 7vw, 64px).
Shape language. Rectilinear and honest. 2px corner radius on cards and buttons — enough to feel built, not rounded-off. 1px and 2px rules do the structural work: a horizontal rule under every page title, a ruled label/value pair for every statistic, a ruled table for every dataset preview. No blobs, no pills, no soft shadows, no continuous-curve radii above 4px. Colour blocks are hard-edged rectangles that butt against each other like signage panels.
Layout. A 12-column grid on a 1280px canvas, collapsing to a single column at 375px. Fixed left rail (240px desktop; a horizontal scrolling chip row on mobile) lists all eleven modules with their numbers 01–11 and a red active marker; the content column is 8 columns wide with a generous 96px top margin. Every page opens with the same masthead: page number, page title, one-line purpose, and a 2px ink rule. Statistics are never cards floating in space — they are ruled rows of label left, numeral right, aligned on a shared baseline. The workflow strip (Data → Cleaning → Analysis → ML → AI → Prediction) is a full-width band of six hard-edged colour panels, each with its stage number reversed out.
Imagery. No stock photography and no 3D. Imagery is diagrammatic: schematic pipeline diagrams, a small set of custom pictograms for each module (database, broom, magnifier, bar chart, encoder, gear, tree, speech, network, spark, crystal ball — drawn on a 24px grid with 2px strokes in ink), Plotly and Matplotlib charts as the primary visual content, and the actual dataset table rendered as first-class imagery. The real confusion matrix and the real loss curve are shown, never mocked up.
Readable-content rule. Headlines, wordmarks, labels, numbers, table cells, chart labels and controls stay entirely inside the viewport and their container at 375px, 768px and 1280px, wrapping or scaling to fit; nothing covers them. Crops and bleeds are for decoration only. The one exception is moving content: the mobile module chip row may scroll horizontally, and every item becomes fully readable as it passes; with prefers-reduced-motion it stops and shows whole items in a horizontally scrollable row.
The Departure Board. The public entry (Home / Dashboard) is not a centred headline with a button. The first screen is a full-width information panel on warm paper #F2EEE5:
clamp(32px, 6vw, 72px), flush left, with the sub-line Data Collection → Cleaning → EDA → Visualization → Preprocessing → Feature Engineering → ML → NLP → Deep Learning → AI/LLM → Prediction set at 17px and wrapping into three ruled rows.Ground is warm paper; no image, no gradient, no glow. The permanent numbered module rail 01–11 runs down the left edge like a transit line: each module a ruled label with its number in mustard, the active one marked by a 4px red bar and a red numeral, connected by a 1px ink rule that fills red as stages are completed. The compact six-panel progress strip repeats at the top of every page so the student always sees where they are in the pipeline. The concept only recomposes accepted content, states and controls — it introduces no new behaviour, page or destination.
Interaction Model: Static (direction) Motion Tempo: restrained Hero Dimensionality: flat
Landing Hero Motion Brief. Focal subject: the ruled statistics stack and the six-panel workflow band on the Home / Dashboard entry. Input → transformation → outcome thesis: as the bundled dataset is read, the four statistics rows resolve their numerals from a placeholder rule into the real record count, feature count, model name and accuracy — the pipeline band's connector rule then fills red left to right, so the first frame becomes a legible departure board of the actual project. Motion vocabulary: functional and short — 160ms ease-out on hover and state change, no bounce, no float, no gradient drift; the one purposeful loop is the pipeline rail, where the active stage marker slides 4px and the connector rule animates its width when a module is completed; charts draw in over 400ms. Composed first frame: headline flush left, three ruled sub-line rows, six colour panels with reversed-out numerals, four ruled statistics rows with 64px tabular numerals, one red PREDICT CAREER button on the headline baseline. Reduced-motion state: with prefers-reduced-motion all of it resolves instantly to the end state — numerals, connector fill and charts appear complete, and the mobile module chip row stops and shows whole items in a horizontally scrollable row.
No user-requested 3D/WebGL and no direction-derived webgl dimensionality, so no Canvas/R3F/Drei requirement applies.
pip install -r requirements.txt followed by streamlit run app.py, and must ship app.py, requirements.txt, README.md and data/career_dataset.csv.prefers-reduced-motion end-state fallback; no gradients, no dark mode, no neon glows or particle fields.Source-specified technology is preserved exactly:
streamlit run app.py.Project files: app.py, requirements.txt, README.md, data/career_dataset.csv.
Constraints (explicit, binding).
app.py, requirements.txt, README.md, data/career_dataset.csv; the website should run using pip install -r requirements.txt and streamlit run app.py.Assumptions (narrow, labeled).
data/career_dataset.csv is the default dataset for every module until the user uploads a replacement on Data Collection; the eight specified columns are present.data/career_dataset.csv containing Education, Specialization, Skills, Programming Level, Experience, Projects, Certifications and Career/Job Role.No completed page designs yet.
Completed design pages will appear here when they are ready to preview.
No completed page designs yet.
Completed design pages will appear here when they are ready to preview.
No comments yet. Be the first!