career-prediction-ai

byregi

Create a simple, clean, fully working **AI & Data Science Predictive Website** for my BCA college project. ## Project Title **AI-Based Career Prediction System** The main purpose of the website is to demonstrate the complete Data Science/AI workflow: **Data Collection → Data Cleaning → EDA → Visualization → Preprocessing → Feature Engineering → Machine Learning → NLP → Deep Learning → AI/LLM → Prediction** The final prediction should use the user's **Education, Skills and Experience** to predict a suitable career/job role. --- # WEBSITE PAGES ## 1. Home / Dashboard Show: * Project title * Short project description * Number of records * Number of features * Machine Learning model used * Model accuracy * Navigation to all modules Also show a simple workflow: **Data → Cleaning → Analysis → ML → AI → Prediction** --- ## 2. Data Collection Purpose: Demonstrate how data is collected. Features: * Upload CSV file * View uploaded dataset * Show number of rows and columns * Show column names * Show first 5/10 records * Allow downloading the dataset Use a sample career dataset containing: * Education * Specialization * Skills * Programming Level * Experience * Projects * Certifications * Career/Job Role --- ## 3. Data Cleaning Purpose: Demonstrate how raw/messy data is cleaned. Show: * Missing values * Duplicate records * Incorrect data types * Missing value handling * Duplicate removal * Cleaned dataset Display before/after statistics. Example: **Before Cleaning** * Rows: 500 * Missing values: 25 * Duplicates: 8 **After Cleaning** * Rows: 492 * Missing values: 0 * Duplicates: 0 --- ## 4. Exploratory Data Analysis (EDA) Purpose: Understand the dataset. Show: * Dataset statistics * Mean * Median * Minimum * Maximum * Standard deviation * Most common education * Most common skill * Most common career Also show useful insights from the dataset. Example: > Python is one of the most common skills among Data Analyst records. --- ## 5. Data Visualization Purpose: Represent data using graphs. Create interactive/simple charts such as: * Education distribution * Skills distribution * Career distribution * Experience distribution * Certification distribution * Education vs Career * Skills vs Career Use: * Matplotlib * Plotly Graphs should update based on the selected dataset. --- ## 6. Data Preprocessing Purpose: Prepare data for Machine Learning. Show: * Categorical encoding * Numerical feature processing * Missing value handling * Feature scaling where required * Train/Test split Explain briefly what preprocessing is doing. Example: **Education → One Hot Encoding** **Experience → Numerical Encoding** **Skills → Multi-label Encoding** --- ## 7. Feature Engineering Purpose: Create useful features from existing data. Create features such as: * Number of Skills * Number of Projects * Experience Score * Certification Score * Programming Skill Score Show the new features in a table. Example: **Python + SQL + ML = Skill Count 3** --- ## 8. Machine Learning Purpose: Train a real Machine Learning model. Use: * Random Forest * Logistic Regression Allow the user to select the model. Show: * Training dataset * Testing dataset * Accuracy * Precision * Recall * F1 Score * Confusion Matrix Display the trained model information. The ML model should actually be used for the final career prediction. --- ## 9. NLP Purpose: Demonstrate Natural Language Processing. Add a small text input: **"Enter your skills or career interest:"** Example: > I know Python, SQL and data visualization. Process the text using: * Text cleaning * Tokenization * TF-IDF Then identify relevant skills/career keywords. Show: **Detected Skills:** * Python * SQL * Data Visualization Also provide a simple sentiment analysis demonstration using sample feedback data. --- ## 10. Deep Learning / Neural Network Purpose: Demonstrate a basic neural network. Create a simple neural network using: * Scikit-learn MLPClassifier or TensorFlow/Keras if appropriate. Use the processed career dataset. Show: * Neural network architecture * Training accuracy * Testing accuracy * Loss/accuracy graph if available Keep this section simple because it is for a BCA academic project. --- ## 11. AI / LLM Purpose: Demonstrate Generative AI. Create a simple **AI Career Assistant**. The user can ask questions such as: > Which skills should I learn for Data Science? > What career can I choose after BCA? > How can I improve my Python skills? Use an LLM API through environment variables. If no API key is available, show a clear message and keep the rest of the website fully functional. Do NOT hardcode any API key. --- # 12. Career Prediction — MAIN PAGE This is the most important page. Create a simple form: ### Education * 10th * 12th * Diploma * BCA * B.Tech * MCA * Other ### Specialization * Computer Science * IT * Data Science * AI/ML * Software Engineering * Other ### Skills Allow multiple selections: * Python * Java * C/C++ * JavaScript * HTML/CSS * SQL * Machine Learning * Data Analysis * Data Visualization * AI * Communication * Problem Solving ### Experience * Fresher * <1 Year * 1–2 Years * 2–5 Years * 5+ Years ### Projects * 0 * 1 * 2–3 * 4+ ### Certification * Yes * No Then add a large button: **🔮 PREDICT CAREER** --- # 13. Prediction Result After prediction, show: ### Predicted Career Example: **Data Analyst** ### Confidence **85%** ### Recommended Skills * Python * SQL * Data Analysis * Power BI ### Why this prediction? Show simple explanations based on the user's input. Example: > Your prediction is influenced by your Python, SQL and Data Analysis skills and your BCA background. Also show the **top 3 possible career roles** with their prediction probabilities, without making the interface complicated. --- # 14. Project Report Create a page containing: * Introduction * Problem Statement * Objectives * Dataset * Data Collection * Data Cleaning * EDA * Visualization * Preprocessing * Feature Engineering * Machine Learning * NLP * Deep Learning * AI/LLM * Prediction * Results * Limitations * Future Scope * Conclusion Add a **Download Report** button. --- # TECHNOLOGY Use: * Python * Streamlit * Pandas * NumPy * Scikit-learn * Matplotlib / Plotly * NLP with TF-IDF * MLP Neural Network * LLM API Keep the project simple and understandable for a BCA student. --- # IMPORTANT REQUIREMENTS The website must be: * Fully working * Simple to understand * Clean and professional * Beginner-friendly * Responsive * No fake buttons * No placeholder pages * No hardcoded prediction results * Real dataset processing * Real Machine Learning prediction * All pages connected through navigation Create these files: ```text app.py requirements.txt README.md data/career_dataset.csv ``` The website should run using: ```bash pip install -r requirements.txt streamlit run app.py ``` Make sure the final project has **one clear purpose: use the complete Data Science/AI workflow to build a Career Prediction System**, while each page demonstrates one specific topic.

No preview

Comments (0)

No comments yet. Be the first!

System Requirements

Page 1 of 21

System Requirements Document for career-prediction-ai

1. Introduction

AI-Based Career Prediction System is a simple, clean, fully working AI & Data Science predictive website built as a BCA college project. Its single purpose is to demonstrate the complete Data Science / AI workflow end to end:

Data Collection → Data Cleaning → EDA → Visualization → Preprocessing → Feature Engineering → Machine Learning → NLP → Deep Learning → AI/LLM → Prediction

Each page of the website demonstrates one specific topic in that pipeline, and the pipeline culminates in a real prediction: the user's Education, Skills and Experience are used to predict a suitable career/job role.

The audience is the BCA student who authors and demonstrates the project, the project guide, and the college examiner. The website must read as real data science rather than a toy: real dataset processing, real Machine Learning prediction, no hardcoded prediction results, no fake buttons, and no placeholder pages. It must remain simple to understand, clean and professional, beginner-friendly, and responsive.

Page 2 of 21

2. System Overview

The system is a single Python/Streamlit application (app.py) that runs locally with:

pip install -r requirements.txt
streamlit run app.py

It ships with four files: app.py, requirements.txt, README.md, and data/career_dataset.csv. The bundled sample career dataset contains the columns Education, Specialization, Skills, Programming Level, Experience, Projects, Certifications, Career/Job Role, and it is the default dataset for every workflow module. A user may also upload their own CSV on the Data Collection page, and every downstream module then operates on that selected dataset.

All fourteen pages are connected through navigation. The left rail lists the eleven numbered pipeline modules (01–11) plus the Home / Dashboard, Career Prediction, Prediction Result, and Project Report destinations, so the student always sees where they are in the pipeline.

Actors:

  • Career Seeker / Student User — enters Education, Specialization, Skills, Experience, Projects and Certification into the Career Prediction form, receives a predicted career with confidence, recommended skills, explanation and top-3 alternatives; asks the AI Career Assistant career questions; describes skills in free text on the NLP page.
  • Project Demonstrator / Evaluator — the student author presenting the project: uploads the sample career dataset, walks through Data Collection, Data Cleaning, EDA, Visualization, Preprocessing, Feature Engineering, Machine Learning, NLP, Deep Learning and AI/LLM to show each Data Science stage working on real data, and uses the Project Report page and its Download Report button to document the work.

Non-persona actors: the LLM API provider (external, accessed only through environment variables) and the application runtime (Streamlit/Python process performing dataset processing, model training and prediction).

Narrow exclusions: no hardcoded API key; no hardcoded prediction results; no fake buttons; no placeholder pages; no blue/indigo SaaS template look; no stock photography, 3D, gradients, glassmorphism or dark mode.

Page 3 of 21

2a. Product Interpretation and Delivery Boundary

Everything in this project is delivered as first-party custom Streamlit UI owned by the application. There is no application account system, no login, no signup, and no per-user stored profile: the source never asks for one, and every accepted journey — uploading a dataset, walking the pipeline, filling the prediction form, reading the result, asking the assistant — is a single-session, anonymous interaction. The Home / Dashboard is the anonymous product entry and every page is reachable without identity.

The only external dependency is the LLM API used by the AI Career Assistant. It is provider-owned and reached exclusively through an environment variable. If no API key is available, the AI / LLM page shows a clear message and the rest of the website remains fully functional. No API key is ever hardcoded.

Current scope is the full eleven-stage pipeline plus the prediction and report destinations. There is no future-horizon section in the source; nothing beyond the accepted pages is committed.

2b. Source Content Inventory

Not applicable — no reference directive declares content_source.

2c. Page Content and Component Coverage

Page 4 of 21

Home / Dashboard

  • Information / state: project title "AI-Based Career Prediction System"; short project description; number of records; number of features; Machine Learning model used; model accuracy; the workflow strip Data → Cleaning → Analysis → ML → AI → Prediction.
  • Primary actions: navigate to all modules (Data Collection, Data Cleaning, EDA, Visualization, Preprocessing, Feature Engineering, Machine Learning, NLP, Deep Learning / Neural Network, AI / LLM, Career Prediction, Prediction Result, Project Report).
  • Supporting actions: jump directly to Career Prediction from the hero.
  • Domain entities: dataset summary (records, features), trained model summary (model name, accuracy).
  • Component responsibilities: masthead with page number, title, one-line purpose and 2px ink rule; hero information panel with headline, pipeline sub-line and ruled statistics stack; six-panel workflow band; numbered module rail 01–11; navigation links to every module.
  • States: loading — statistics resolve while the bundled dataset is read; empty — if no dataset is available, statistics rows show an explicit unavailable state instead of invented numbers; success — records, features, model name and accuracy displayed as ruled label/numeral rows; error — dataset read failure shows a clear message and the navigation remains usable; recovery — the user can open Data Collection to upload a dataset and return.

Data Collection

  • Information / state: uploaded dataset preview; number of rows and columns; column names; first 5/10 records; the bundled sample career dataset with Education, Specialization, Skills, Programming Level, Experience, Projects, Certifications, Career/Job Role.
  • Primary actions: upload a CSV file; view the uploaded dataset; download the dataset.
  • Supporting actions: choose how many leading records to display (5 or 10); switch back to the bundled sample dataset.
  • Domain entities: dataset, column, record.
  • Component responsibilities: file uploader; dataset table ruled like a timetable; row/column counters; column-name list; record-count selector; download control.
  • States: loading — upload in progress; empty — no file uploaded yet, sample dataset shown as the default; success — dataset rendered with rows, columns, column names and first records; error — unreadable or malformed CSV shows a clear message and keeps the previously selected dataset; recovery — re-upload or fall back to the sample dataset.

Data Cleaning

  • Information / state: missing values; duplicate records; incorrect data types; missing value handling; duplicate removal; cleaned dataset; before/after statistics (rows, missing values, duplicates).
  • Primary actions: run cleaning on the selected dataset; view the cleaned dataset.
  • Supporting actions: inspect which columns carry missing values, duplicates or wrong types.
  • Domain entities: raw dataset, cleaned dataset, missing-value count, duplicate count, column dtype.
  • Component responsibilities: before/after statistics panel (e.g. Before: Rows 500, Missing values 25, Duplicates 8 → After: Rows 492, Missing values 0, Duplicates 0); issue list per column; cleaned-dataset table.
  • States: loading — cleaning in progress; empty — no dataset selected, prompt to visit Data Collection; success — before/after statistics and cleaned dataset displayed; error — cleaning failure on an incompatible column shows the offending column and leaves the raw dataset intact; recovery — fix the source CSV or re-run cleaning.
Page 5 of 21

Exploratory Data Analysis (EDA)

  • Information / state: dataset statistics — mean, median, minimum, maximum, standard deviation; most common education; most common skill; most common career; useful dataset insights (e.g. "Python is one of the most common skills among Data Analyst records").
  • Primary actions: compute and view statistics for the selected dataset.
  • Supporting actions: read the generated insight statements.
  • Domain entities: numeric column statistics, categorical frequency counts, insight statement.
  • Component responsibilities: ruled statistics rows; most-common value panels; insight list.
  • States: loading — statistics computing; empty — no dataset selected; success — statistics and insights rendered; error — non-numeric columns excluded from mean/median/min/max/std with a note rather than a crash; recovery — select a different dataset or column.

Data Visualization

  • Information / state: charts for Education distribution, Skills distribution, Career distribution, Experience distribution, Certification distribution, Education vs Career, Skills vs Career.
  • Primary actions: render charts for the selected dataset.
  • Supporting actions: hover/zoom on interactive Plotly charts.
  • Domain entities: distribution counts, cross-tabulations.
  • Component responsibilities: Matplotlib figures and Plotly interactive charts drawn on the paper ground with a 1px ink baseline and no gridlines; chart selector.
  • States: loading — charts drawing in over 400ms; empty — no dataset selected; success — all seven chart types render and update when the selected dataset changes; error — a chart that cannot be built for the current columns shows a clear message while the others still render; recovery — pick another chart or dataset.

Data Preprocessing

  • Information / state: categorical encoding; numerical feature processing; missing value handling; feature scaling where required; train/test split; brief explanations of what preprocessing is doing (e.g. Education → One Hot Encoding, Experience → Numerical Encoding, Skills → Multi-label Encoding).
  • Primary actions: run preprocessing on the selected dataset.
  • Supporting actions: read the per-step explanation text.
  • Domain entities: encoded feature matrix, scaled numeric features, train split, test split.
  • Component responsibilities: step-by-step explanation blocks; encoding summary; train/test split summary with sizes.
  • States: loading — preprocessing running; empty — no dataset selected; success — encodings, scaling and split sizes displayed; error — an unencodable column is reported with its name; recovery — adjust the dataset or re-run.
Page 6 of 21

Feature Engineering

  • Information / state: new features — Number of Skills, Number of Projects, Experience Score, Certification Score, Programming Skill Score — shown in a table (e.g. Python + SQL + ML = Skill Count 3).
  • Primary actions: create the engineered features for the selected dataset.
  • Supporting actions: inspect the new-feature table row by row.
  • Domain entities: engineered feature, source columns used.
  • Component responsibilities: engineered-feature table; per-feature derivation note.
  • States: loading — features being derived; empty — no dataset selected; success — new features table rendered; error — a missing source column is named explicitly; recovery — select a dataset containing the required columns.

Machine Learning

  • Information / state: training dataset; testing dataset; accuracy; precision; recall; F1 score; confusion matrix; trained model information.
  • Primary actions: select the model (Random Forest or Logistic Regression); train the model.
  • Supporting actions: inspect the confusion matrix and the trained model information.
  • Domain entities: model type, training set, test set, evaluation metrics, confusion matrix, trained model.
  • Component responsibilities: model selector; metric rows using the departure-board treatment; confusion matrix display; trained-model information panel.
  • States: loading — training in progress; empty — no processed dataset available, prompt to complete preprocessing; success — metrics and confusion matrix displayed for the selected model; error — training failure shows a clear message; recovery — switch model or re-run preprocessing.
  • Note: the trained model is the model actually used for the final career prediction.

NLP

  • Information / state: text input labelled "Enter your skills or career interest:"; processed text; Detected Skills list (e.g. Python, SQL, Data Visualization); sentiment analysis demonstration using sample feedback data.
  • Primary actions: submit free text for processing; view detected skills; view the sentiment demonstration.
  • Supporting actions: read the text-cleaning, tokenization and TF-IDF steps applied.
  • Domain entities: raw text, cleaned text, tokens, TF-IDF vectors, detected skill/career keywords, sentiment result.
  • Component responsibilities: text input; processing-step display; detected-skills list; sentiment demonstration panel over sample feedback data.
  • States: loading — text processing; empty — no text entered yet; success — detected skills and sentiment shown; error — empty or unprocessable input shows a clear prompt; recovery — edit the text and resubmit.
Page 7 of 21

Deep Learning / Neural Network

  • Information / state: neural network architecture; training accuracy; testing accuracy; loss/accuracy graph if available.
  • Primary actions: train the simple neural network on the processed career dataset.
  • Supporting actions: inspect the architecture summary and the loss/accuracy graph.
  • Domain entities: MLPClassifier (or TensorFlow/Keras if appropriate), architecture layers, training accuracy, testing accuracy, loss curve.
  • Component responsibilities: architecture display; accuracy rows; loss/accuracy chart.
  • States: loading — network training; empty — no processed dataset available; success — architecture, accuracies and graph displayed; error — training failure shows a clear message; recovery — re-run after preprocessing.
  • Note: kept deliberately simple because it is for a BCA academic project.

AI / LLM

  • Information / state: AI Career Assistant question input and answer area; API-availability state.
  • Primary actions: ask a career question (e.g. "Which skills should I learn for Data Science?", "What career can I choose after BCA?", "How can I improve my Python skills?"); read the generated answer.
  • Supporting actions: none beyond asking and reading.
  • Domain entities: question, generated answer, API key availability.
  • Component responsibilities: question input; answer panel; unavailable-key message.
  • States: loading — answer being generated; empty — no question asked yet; success — generated answer displayed; error — when no API key is available, a clear message is shown and the rest of the website stays fully functional; recovery — set the environment variable and retry, or continue using the other modules.
  • Constraint: the LLM API is accessed through environment variables only; no API key is hardcoded.
Page 8 of 21

Career Prediction

  • Information / state: the main prediction form with Education (10th, 12th, Diploma, BCA, B.Tech, MCA, Other), Specialization (Computer Science, IT, Data Science, AI/ML, Software Engineering, Other), multi-select Skills (Python, Java, C/C++, JavaScript, HTML/CSS, SQL, Machine Learning, Data Analysis, Data Visualization, AI, Communication, Problem Solving), Experience (Fresher, <1 Year, 1–2 Years, 2–5 Years, 5+ Years), Projects (0, 1, 2–3, 4+), Certification (Yes, No).
  • Primary actions: complete the form; press the large \xf0\x9f\x94\xae PREDICT CAREER button.
  • Supporting actions: change any field and predict again.
  • Domain entities: education, specialization, skills, experience, projects, certification, prediction request.
  • Component responsibilities: form controls for each field; the large predict button; validation feedback.
  • States: loading — prediction running against the trained model; empty — form not yet submitted; success — the Prediction Result is produced from the real trained model; error — incomplete required input shows a clear validation message; recovery — complete the missing field and predict again.
  • Constraint: no hardcoded prediction results; the ML model is actually used.

Prediction Result

  • Information / state: predicted career (e.g. Data Analyst); confidence percentage (e.g. 85%); recommended skills (e.g. Python, SQL, Data Analysis, Power BI); "Why this prediction?" simple explanations based on the user's input (e.g. "Your prediction is influenced by your Python, SQL and Data Analysis skills and your BCA background"); the top 3 possible career roles with their prediction probabilities.
  • Primary actions: read the prediction, confidence, recommendations, explanation and top-3 alternatives.
  • Supporting actions: return to Career Prediction to adjust inputs and predict again.
  • Domain entities: predicted career, confidence, recommended skills, explanation, top-3 roles with probabilities.
  • Component responsibilities: signage panel with the predicted career at 44px ink; bare 64px red confidence numeral with no progress bar; recommended-skills list; explanation block; three ruled rows for the top-3 roles with probabilities right-aligned.
  • States: loading — result being assembled; empty — no prediction has been made yet, prompt to open Career Prediction; success — full result displayed without complicating the interface; error — prediction failure shows a clear message; recovery — re-submit the form.

Project Report

  • Information / state: Introduction, Problem Statement, Objectives, Dataset, Data Collection, Data Cleaning, EDA, Visualization, Preprocessing, Feature Engineering, Machine Learning, NLP, Deep Learning, AI/LLM, Prediction, Results, Limitations, Future Scope, Conclusion.
  • Primary actions: read the report; press the Download Report button.
  • Supporting actions: navigate between report sections.
  • Domain entities: report section, report document.
  • Component responsibilities: sectioned report body; download control producing the report file.
  • States: loading — report assembling; empty — not applicable, the report content is always present; success — all sections rendered and the download produces a real file; error — download failure shows a clear message; recovery — retry the download.
Page 9 of 21

3. Functional Requirements

FR-1 — Pipeline demonstration (explicit). As a Project Demonstrator / Evaluator, I should be able to walk the complete workflow Data Collection → Data Cleaning → EDA → Visualization → Preprocessing → Feature Engineering → Machine Learning → NLP → Deep Learning → AI/LLM → Prediction across the website, so that each page demonstrates one specific Data Science/AI topic. Lifecycle: initiator = demonstrator; trigger = opening the app; observable result = each stage page renders its own real output; failure/recovery = a stage that cannot run names the missing prerequisite and links back to it; continuation = the next stage in the pipeline. Acceptance: all eleven stages are reachable and each produces real output from the selected dataset.

FR-2 — Home / Dashboard (explicit). As a Career Seeker / Student User or Project Demonstrator / Evaluator, I should see the project title, a short project description, the number of records, the number of features, the Machine Learning model used, the model accuracy, navigation to all modules, and the workflow Data → Cleaning → Analysis → ML → AI → Prediction. Lifecycle: initiator = visitor; trigger = opening the app; observable result = dashboard statistics and workflow strip; failure/recovery = unavailable statistics are labelled rather than invented; continuation = navigate to any module. Acceptance: every listed item is visible and every module is reachable.

FR-3 — Data Collection (explicit). As a Project Demonstrator / Evaluator, I should upload a CSV file, view the uploaded dataset, see the number of rows and columns, see the column names, see the first 5/10 records, and download the dataset. Lifecycle: initiator = demonstrator; trigger = CSV upload or default sample load; observable result = dataset preview with row/column counts, column names and leading records; failure/recovery = malformed CSV shows a clear message and keeps the previous dataset; continuation = proceed to Data Cleaning. Acceptance: the bundled sample career dataset contains Education, Specialization, Skills, Programming Level, Experience, Projects, Certifications, Career/Job Role, and the download produces the selected dataset.

FR-4 — Data Cleaning (explicit). As a Project Demonstrator / Evaluator, I should see missing values, duplicate records and incorrect data types, apply missing value handling and duplicate removal, view the cleaned dataset, and see before/after statistics. Lifecycle: initiator = demonstrator; trigger = run cleaning; observable result = before/after rows, missing values and duplicates (e.g. Before: Rows 500, Missing values 25, Duplicates 8 → After: Rows 492, Missing values 0, Duplicates 0) plus the cleaned dataset; failure/recovery = an incompatible column is named and the raw dataset is preserved; continuation = proceed to EDA. Acceptance: before/after statistics and the cleaned dataset are both displayed.

FR-5 — Exploratory Data Analysis (explicit). As a Project Demonstrator / Evaluator, I should see dataset statistics (mean, median, minimum, maximum, standard deviation), the most common education, the most common skill, the most common career, and useful dataset insights. Lifecycle: initiator = demonstrator; trigger = open EDA with a selected dataset; observable result = statistics, most-common values and insight statements such as "Python is one of the most common skills among Data Analyst records"; failure/recovery = non-numeric columns are excluded from numeric statistics with a note; continuation = proceed to Visualization. Acceptance: all listed statistics and the three most-common values are shown.

FR-6 — Data Visualization (explicit). As a Project Demonstrator / Evaluator, I should see charts for Education distribution, Skills distribution, Career distribution, Experience distribution, Certification distribution, Education vs Career, and Skills vs Career, built with Matplotlib and Plotly, updating based on the selected dataset. Lifecycle: initiator = demonstrator; trigger = open Visualization or change the selected dataset; observable result = the seven chart types re-render for the current dataset; failure/recovery = a chart that cannot be built is reported while the others still render; continuation = proceed to Preprocessing. Acceptance: all seven chart types exist and update when the selected dataset changes.

FR-7 — Data Preprocessing (explicit). As a Project Demonstrator / Evaluator, I should see categorical encoding, numerical feature processing, missing value handling, feature scaling where required, and the train/test split, with brief explanations of what preprocessing is doing (e.g. Education → One Hot Encoding, Experience → Numerical Encoding, Skills → Multi-label Encoding). Lifecycle: initiator = demonstrator; trigger = run preprocessing; observable result = encoding summary, scaling note and train/test split sizes; failure/recovery = an unencodable column is named; continuation = proceed to Feature Engineering. Acceptance: each preprocessing step is shown with its explanation.

FR-8 — Feature Engineering (explicit). As a Project Demonstrator / Evaluator, I should create Number of Skills, Number of Projects, Experience Score, Certification Score and Programming Skill Score, and see the new features in a table. Lifecycle: initiator = demonstrator; trigger = run feature engineering; observable result = engineered-feature table (e.g. Python + SQL + ML = Skill Count 3); failure/recovery = a missing source column is named explicitly; continuation = proceed to Machine Learning. Acceptance: all five engineered features appear in the table.

FR-9 — Machine Learning (explicit). As a Project Demonstrator / Evaluator, I should select Random Forest or Logistic Regression, train a real model, and see the training dataset, testing dataset, accuracy, precision, recall, F1 score, confusion matrix and trained model information. Lifecycle: initiator = demonstrator; trigger = model selection and training; observable result = metrics and confusion matrix for the selected model; failure/recovery = training failure shows a clear message and the model can be switched; continuation = the trained model is used for the final career prediction. Acceptance: both models are selectable, all six evaluation outputs are shown, and the trained model actually drives prediction.

FR-10 — NLP (explicit). As a Career Seeker / Student User, I should type into the input "Enter your skills or career interest:" (e.g. "I know Python, SQL and data visualization."), have the text processed with text cleaning, tokenization and TF-IDF, and see the relevant skills/career keywords under Detected Skills (e.g. Python, SQL, Data Visualization). Lifecycle: initiator = student; trigger = submitting free text; observable result = detected skills list; failure/recovery = empty or unprocessable input shows a clear prompt; continuation = use the detected skills to inform the Career Prediction form. Acceptance: the three processing steps are applied and the detected skills are displayed.

FR-11 — NLP sentiment demonstration (explicit). As a Project Demonstrator / Evaluator, I should see a simple sentiment analysis demonstration using sample feedback data. Lifecycle: initiator = demonstrator; trigger = open the sentiment demonstration; observable result = sentiment result over the sample feedback data; failure/recovery = unavailable sample feedback shows a clear message; continuation = return to the NLP keyword flow. Acceptance: the sentiment demonstration renders on the NLP page.

FR-12 — Deep Learning / Neural Network (explicit). As a Project Demonstrator / Evaluator, I should train a simple neural network using Scikit-learn MLPClassifier (or TensorFlow/Keras if appropriate) on the processed career dataset and see the neural network architecture, training accuracy, testing accuracy, and a loss/accuracy graph if available. Lifecycle: initiator = demonstrator; trigger = train the network; observable result = architecture, both accuracies and the graph; failure/recovery = training failure shows a clear message; continuation = compare with the Machine Learning page results. Acceptance: the section stays simple for a BCA academic project and shows all four listed outputs.

FR-13 — AI Career Assistant (explicit). As a Career Seeker / Student User, I should ask questions such as "Which skills should I learn for Data Science?", "What career can I choose after BCA?" and "How can I improve my Python skills?" and receive answers from an LLM API accessed through environment variables. Lifecycle: initiator = student; trigger = submitting a question; observable result = generated answer; failure/recovery = if no API key is available, a clear message is shown and the rest of the website remains fully functional; continuation = ask another question or continue to Career Prediction. Acceptance: no API key is hardcoded anywhere, and the unavailable-key state is explicit.

FR-14 — Career Prediction form (explicit). As a Career Seeker / Student User, I should fill a simple form with Education (10th, 12th, Diploma, BCA, B.Tech, MCA, Other), Specialization (Computer Science, IT, Data Science, AI/ML, Software Engineering, Other), multi-select Skills (Python, Java, C/C++, JavaScript, HTML/CSS, SQL, Machine Learning, Data Analysis, Data Visualization, AI, Communication, Problem Solving), Experience (Fresher, <1 Year, 1–2 Years, 2–5 Years, 5+ Years), Projects (0, 1, 2–3, 4+) and Certification (Yes, No), then press the large \xf0\x9f\x94\xae PREDICT CAREER button. Lifecycle: initiator = student; trigger = completing the form and pressing the button; observable result = a prediction request is executed against the real trained model; failure/recovery = incomplete required input shows a clear validation message; continuation = the Prediction Result is shown. Acceptance: every listed option is present exactly as specified and the button is large and functional.

FR-15 — Prediction Result (explicit). As a Career Seeker / Student User, I should see the predicted career (e.g. Data Analyst), the confidence (e.g. 85%), recommended skills (e.g. Python, SQL, Data Analysis, Power BI), a "Why this prediction?" explanation based on my input (e.g. "Your prediction is influenced by your Python, SQL and Data Analysis skills and your BCA background"), and the top 3 possible career roles with their prediction probabilities, without the interface becoming complicated. Lifecycle: initiator = student; trigger = a completed prediction; observable result = the full result panel; failure/recovery = prediction failure shows a clear message and the form can be resubmitted; continuation = adjust inputs and predict again. Acceptance: all five result elements are shown and no result is hardcoded.

FR-16 — Project Report (explicit). As a Project Demonstrator / Evaluator, I should read a report containing Introduction, Problem Statement, Objectives, Dataset, Data Collection, Data Cleaning, EDA, Visualization, Preprocessing, Feature Engineering, Machine Learning, NLP, Deep Learning, AI/LLM, Prediction, Results, Limitations, Future Scope and Conclusion, and press a Download Report button. Lifecycle: initiator = demonstrator; trigger = opening the report or pressing download; observable result = the report content and a downloaded report file; failure/recovery = download failure shows a clear message; continuation = return to any module. Acceptance: every listed section is present and the download button produces a real file.

FR-17 — Sample dataset availability (required_inference). As a Project Demonstrator / Evaluator, I should have the sample career dataset available at data/career_dataset.csv for the workflow modules and real model prediction. Lifecycle: initiator = application runtime on first load; trigger = app start; observable result = the sample dataset is the default selection for every module; failure/recovery = a missing file shows a clear message and prompts upload on Data Collection; continuation = the pipeline runs on the uploaded dataset. Acceptance: the bundled dataset contains the eight specified columns.

FR-18 — Dataset processing prerequisite (required_inference). As a Project Demonstrator / Evaluator, I should have a submitted or sample dataset processed before dependent statistics, charts, preprocessing, feature engineering, model evaluation, neural-network and prediction views can operate. Lifecycle: initiator = application runtime; trigger = opening a dependent module; observable result = dependent modules render from the currently selected dataset; failure/recovery = a module without a usable dataset shows a clear prerequisite message and a route to Data Collection; continuation = the module renders once a dataset is selected. Acceptance: no dependent module silently fabricates data.

FR-19 — LLM key prerequisite (required_inference). As a Career Seeker / Student User, I should be able to use the AI Career Assistant only when an LLM API key is supplied through an environment variable; without it the assistant shows its unavailable-key message while the other modules remain functional. Lifecycle: initiator = student; trigger = asking a question; observable result = answer or unavailable-key message; failure/recovery = the message names the missing environment variable; continuation = the rest of the website stays usable. Acceptance: the key is read from the environment only.

FR-20 — Navigation connectivity (explicit). As a Career Seeker / Student User or Project Demonstrator / Evaluator, I should reach every page through navigation, with no fake buttons and no placeholder pages. Lifecycle: initiator = any visitor; trigger = using the module rail or in-page links; observable result = the target page opens with real content; failure/recovery = a link to an unavailable state shows the reason; continuation = the user returns to the pipeline. Acceptance: all fourteen pages are connected and every control performs a real action.

Page 10 of 21

4. User Personas

Page 11 of 21

Career Seeker / Student User

Provenance: required_inference (Planning Scope).

Product context. A BCA student or job seeker who wants to know which career role fits their Education, Specialization, Skills, Experience, Projects and Certification. They are the reason the prediction exists; the pipeline pages are how they come to trust the answer.

Primary goal. Enter their own profile once and receive a suitable career/job role with a confidence figure, recommended skills, a plain-language explanation of why, and the top 3 alternative roles.

Distinct accepted responsibilities. Completing the Career Prediction form with the exact Education, Specialization, multi-select Skills, Experience, Projects and Certification options; pressing \xf0\x9f\x94\xae PREDICT CAREER; reading the predicted career, confidence, recommended skills, "Why this prediction?" explanation and top-3 alternatives; describing their skills in free text on the NLP page ("I know Python, SQL and data visualization.") and reading the Detected Skills; asking the AI Career Assistant questions such as "Which skills should I learn for Data Science?", "What career can I choose after BCA?" and "How can I improve my Python skills?".

Relevant inputs or decisions. Which education level and specialization describe them; which of the twelve skills they actually have; their experience band; how many projects they have completed; whether they hold a certification; whether to accept the predicted role or adjust their inputs and predict again.

Interactions with other accepted participants. They consume the model that the Project Demonstrator / Evaluator trained on the Machine Learning page; the AI Career Assistant answers come from the external LLM API provider when a key is configured.

Observable success. A prediction result panel showing a predicted career, a confidence percentage, recommended skills, an explanation referencing their own inputs, and three ranked alternative roles — produced by the real trained model, never hardcoded.

Page 12 of 21

Project Demonstrator / Evaluator

Provenance: required_inference (Planning Scope).

Product context. The student author presenting the capstone to a project guide or examiner, and the evaluator assessing whether the whole Data Science/AI workflow genuinely runs. Their credibility depends on every stage showing real output from real data.

Primary goal. Demonstrate a fully working, connected pipeline in which each page proves one Data Science/AI topic on the actual dataset, and hand over a documented report.

Distinct accepted responsibilities. Uploading the sample career dataset and inspecting rows, columns, column names and the first 5/10 records; downloading the dataset; running cleaning and showing before/after statistics; reading EDA statistics and insights; rendering the seven visualization chart types; running preprocessing and explaining the encodings; creating and showing the engineered features; selecting Random Forest or Logistic Regression and presenting accuracy, precision, recall, F1 and the confusion matrix; running the NLP keyword and sentiment demonstrations; training the MLP neural network and showing architecture, training accuracy, testing accuracy and the loss/accuracy graph; demonstrating the AI Career Assistant and its unavailable-key behaviour; and presenting the Project Report with its Download Report button.

Relevant inputs or decisions. Which dataset to demonstrate with; which model to select; which chart or statistic to highlight; whether the LLM key is configured for the live demonstration.

Interactions with other accepted participants. They produce the trained model and the report that the Career Seeker / Student User depends on; they present the AI / LLM page whose answers come from the external LLM API provider.

Observable success. Every module renders real output with no placeholder pages and no fake buttons, the pipeline is navigable end to end, and the report downloads as a real file.

Page 13 of 21

5. Core User Flows

Flow A — Career Seeker / Student User: get a career prediction

  1. The student opens the app and lands on Home / Dashboard, which shows the project title, description, number of records, number of features, the Machine Learning model used, the model accuracy, and the workflow Data → Cleaning → Analysis → ML → AI → Prediction.
  2. The student selects Career Prediction from the module rail.
  3. On Career Prediction, the student sets Education (e.g. BCA), Specialization (e.g. Computer Science), multi-selects Skills (e.g. Python, SQL, Data Analysis), sets Experience (e.g. Fresher), Projects (e.g. 2–3) and Certification (e.g. Yes).
  4. The student presses the large \xf0\x9f\x94\xae PREDICT CAREER button. If a required field is missing, a clear validation message names it and the student completes it and presses the button again.
  5. The application runs the completed profile through the real trained model and opens Prediction Result.
  6. Prediction Result shows the predicted career (e.g. Data Analyst), the confidence (e.g. 85%), recommended skills (e.g. Python, SQL, Data Analysis, Power BI), the "Why this prediction?" explanation referencing the student's own inputs, and the top 3 possible career roles with their probabilities.
  7. Continuation: the student returns to Career Prediction, changes a field (for example adds Machine Learning to Skills) and predicts again to compare results.

Flow B — Career Seeker / Student User: describe skills in free text

  1. The student opens NLP.
  2. In the input labelled "Enter your skills or career interest:" the student types, for example, "I know Python, SQL and data visualization."
  3. The application applies text cleaning, tokenization and TF-IDF.
  4. Detected Skills is displayed: Python, SQL, Data Visualization.
  5. Continuation: the student carries those detected skills into the Career Prediction form. If the input is empty or unprocessable, a clear prompt asks for text and the student resubmits.

Flow C — Career Seeker / Student User: ask the AI Career Assistant

  1. The student opens AI / LLM.
  2. The student asks, for example, "Which skills should I learn for Data Science?", "What career can I choose after BCA?" or "How can I improve my Python skills?".
  3. The application reads the LLM API key from the environment variable and sends the question to the LLM API provider.
  4. The generated answer is displayed on the page.
  5. Failure/recovery: if no API key is available, the page shows a clear message naming the missing environment variable, and every other module of the website remains fully functional.
  6. Continuation: the student asks another question, or moves to Career Prediction.
Page 14 of 21

Flow D — Project Demonstrator / Evaluator: demonstrate the pipeline

  1. The demonstrator opens Home / Dashboard and points out the records, features, model and accuracy figures and the workflow strip.
  2. On Data Collection, the demonstrator uploads a CSV or uses the bundled sample career dataset, shows the number of rows and columns, the column names, the first 5/10 records, and downloads the dataset.
  3. On Data Cleaning, the demonstrator runs cleaning and shows the before/after statistics (e.g. Before: Rows 500, Missing values 25, Duplicates 8 → After: Rows 492, Missing values 0, Duplicates 0) and the cleaned dataset.
  4. On Exploratory Data Analysis (EDA), the demonstrator shows mean, median, minimum, maximum, standard deviation, the most common education, skill and career, and the generated insights.
  5. On Data Visualization, the demonstrator renders the Education, Skills, Career, Experience and Certification distributions plus Education vs Career and Skills vs Career, and changes the selected dataset to show the charts updating.
  6. On Data Preprocessing, the demonstrator shows categorical encoding, numerical feature processing, missing value handling, feature scaling where required and the train/test split, with the brief explanations (Education → One Hot Encoding, Experience → Numerical Encoding, Skills → Multi-label Encoding).
  7. On Feature Engineering, the demonstrator creates Number of Skills, Number of Projects, Experience Score, Certification Score and Programming Skill Score and shows them in the new-features table.
  8. On Machine Learning, the demonstrator selects Random Forest (then Logistic Regression), trains, and shows the training dataset, testing dataset, accuracy, precision, recall, F1 score, confusion matrix and trained model information.
  9. On NLP, the demonstrator shows the text-cleaning, tokenization and TF-IDF steps, the Detected Skills output, and the sentiment analysis demonstration over sample feedback data.
  10. On Deep Learning / Neural Network, the demonstrator trains the simple MLPClassifier on the processed career dataset and shows the architecture, training accuracy, testing accuracy and the loss/accuracy graph.
  11. On AI / LLM, the demonstrator asks a sample career question and shows either the generated answer or the clear unavailable-key message.
  12. On Career Prediction, the demonstrator fills the form and presses \xf0\x9f\x94\xae PREDICT CAREER, then shows the Prediction Result produced by the model trained in step 8.
  13. On Project Report, the demonstrator walks the Introduction, Problem Statement, Objectives, Dataset, Data Collection, Data Cleaning, EDA, Visualization, Preprocessing, Feature Engineering, Machine Learning, NLP, Deep Learning, AI/LLM, Prediction, Results, Limitations, Future Scope and Conclusion sections, and presses Download Report to produce the report file.
  14. Failure/recovery: if any stage cannot run because no dataset is selected, that stage shows a clear prerequisite message and a route back to Data Collection; the demonstrator selects a dataset and continues from that stage.
Page 15 of 21

6. Visuals Colors and Theme

The creative direction is authoritative: typographic infrastructure for a career pipeline, after the muse Erik Spiekermann — typography as infrastructure, transit signage, numbered modules, data read like a departure board. Headline: an eleven-stage pipeline that reads like a public information system, warm enough for a classroom.

Colour tokens (light mode only — no dark mode).

RoleHexUse
Background#F2EEE5Warm paper ground for the whole page
Surface#FFFDF7Cards and data panels, with a 1px #D8D1C2 rule, never a shadow
Text / ink#1A1A18All reading text and numerals
Primary#C7401FTransit red: active module in the pipeline rail, PREDICT CAREER button, predicted-career headline, accuracy figures
Accent#E8B22BMustard: secondary line colour for the workflow strip, chart series and stage numbers
Reserved line#1F6E63Deep teal, used only for the neural network and LLM stages so the pipeline reads like separate transit lines
Muted#6F6A5ECaptions, metadata and axis labels
Rule#D8D1C21px structural rules, table row separators

Charts use the four signal colours at full saturation on the paper ground. No gradients anywhere. Blue and indigo are forbidden (#0057FF, #2563EB, #4F46E5, #6366F1 and neighbours), as is the blue-on-white default Streamlit theme.

Typography. Headings and body: Fira Sans (Spiekermann's own humanist sans). Headings at 700–800 weight, tight -0.02em tracking, sentence case for page titles, all-caps with +0.14em tracking for module numbers, stage labels and table headers. Numerals are the ornament: model accuracy, record counts and confidence percentages are set large in 700 weight, tabular, so they read like departure-board figures. No italics, no serifs, no display face. Forbidden for headings or body: Inter, Roboto, Arial, Helvetica, Open Sans, Lato, Poppins, system-ui.

Type scale (1.25 modular). Mobile 30/24/19/16; desktop 44/32/22/17. Hero title clamp(32px, 6vw, 72px); page titles clamp(30px, 4.5vw, 44px); section titles 22px; body 17px/1.55; labels 13px all-caps; numerals clamp(36px, 7vw, 64px).

Shape language. Rectilinear and honest. 2px corner radius on cards and buttons — enough to feel built, not rounded-off. 1px and 2px rules do the structural work: a horizontal rule under every page title, a ruled label/value pair for every statistic, a ruled table for every dataset preview. No blobs, no pills, no soft shadows, no continuous-curve radii above 4px. Colour blocks are hard-edged rectangles that butt against each other like signage panels.

Layout. A 12-column grid on a 1280px canvas, collapsing to a single column at 375px. Fixed left rail (240px desktop; a horizontal scrolling chip row on mobile) lists all eleven modules with their numbers 01–11 and a red active marker; the content column is 8 columns wide with a generous 96px top margin. Every page opens with the same masthead: page number, page title, one-line purpose, and a 2px ink rule. Statistics are never cards floating in space — they are ruled rows of label left, numeral right, aligned on a shared baseline. The workflow strip (Data → Cleaning → Analysis → ML → AI → Prediction) is a full-width band of six hard-edged colour panels, each with its stage number reversed out.

Imagery. No stock photography and no 3D. Imagery is diagrammatic: schematic pipeline diagrams, a small set of custom pictograms for each module (database, broom, magnifier, bar chart, encoder, gear, tree, speech, network, spark, crystal ball — drawn on a 24px grid with 2px strokes in ink), Plotly and Matplotlib charts as the primary visual content, and the actual dataset table rendered as first-class imagery. The real confusion matrix and the real loss curve are shown, never mocked up.

Readable-content rule. Headlines, wordmarks, labels, numbers, table cells, chart labels and controls stay entirely inside the viewport and their container at 375px, 768px and 1280px, wrapping or scaling to fit; nothing covers them. Crops and bleeds are for decoration only. The one exception is moving content: the mobile module chip row may scroll horizontally, and every item becomes fully readable as it passes; with prefers-reduced-motion it stops and shows whole items in a horizontally scrollable row.

Page 16 of 21

7. Signature Design Concept

The Departure Board. The public entry (Home / Dashboard) is not a centred headline with a button. The first screen is a full-width information panel on warm paper #F2EEE5:

  • Left: a three-line Fira Sans 800 headline AI-BASED CAREER PREDICTION SYSTEM at clamp(32px, 6vw, 72px), flush left, with the sub-line Data Collection → Cleaning → EDA → Visualization → Preprocessing → Feature Engineering → ML → NLP → Deep Learning → AI/LLM → Prediction set at 17px and wrapping into three ruled rows.
  • Beneath it: a horizontal band of six hard-edged colour panels — red, mustard, teal, ink, red, mustard — reading DATA / CLEANING / ANALYSIS / ML / AI / PREDICTION, each with its numeral reversed out. The band bleeds to the right edge as decoration only; the panel text itself stays whole.
  • Right: a ruled statistics stack on the paper ground — 500 RECORDS, 8 FEATURES, RANDOM FOREST, 92% ACCURACY as four label-left / numeral-right rows with 1px rules between them, the numerals at 64px tabular.
  • Baseline: a single red PREDICT CAREER button, square-cornered at 2px radius, sitting on the baseline of the headline block.

Ground is warm paper; no image, no gradient, no glow. The permanent numbered module rail 01–11 runs down the left edge like a transit line: each module a ruled label with its number in mustard, the active one marked by a 4px red bar and a red numeral, connected by a 1px ink rule that fills red as stages are completed. The compact six-panel progress strip repeats at the top of every page so the student always sees where they are in the pipeline. The concept only recomposes accepted content, states and controls — it introduces no new behaviour, page or destination.

Page 17 of 21

8. Interaction Model & Motion Direction

Interaction Model: Static (direction) Motion Tempo: restrained Hero Dimensionality: flat

Landing Hero Motion Brief. Focal subject: the ruled statistics stack and the six-panel workflow band on the Home / Dashboard entry. Input → transformation → outcome thesis: as the bundled dataset is read, the four statistics rows resolve their numerals from a placeholder rule into the real record count, feature count, model name and accuracy — the pipeline band's connector rule then fills red left to right, so the first frame becomes a legible departure board of the actual project. Motion vocabulary: functional and short — 160ms ease-out on hover and state change, no bounce, no float, no gradient drift; the one purposeful loop is the pipeline rail, where the active stage marker slides 4px and the connector rule animates its width when a module is completed; charts draw in over 400ms. Composed first frame: headline flush left, three ruled sub-line rows, six colour panels with reversed-out numerals, four ruled statistics rows with 64px tabular numerals, one red PREDICT CAREER button on the headline baseline. Reduced-motion state: with prefers-reduced-motion all of it resolves instantly to the end state — numerals, connector fill and charts appear complete, and the mobile module chip row stops and shows whole items in a horizontally scrollable row.

No user-requested 3D/WebGL and no direction-derived webgl dimensionality, so no Canvas/R3F/Drei requirement applies.

Page 18 of 21

9. Non-Functional Requirements

  • NFR-1 (explicit). The website must be fully working: every control performs a real action, with no fake buttons and no placeholder pages.
  • NFR-2 (explicit). Real dataset processing and real Machine Learning prediction are required; prediction results must never be hardcoded.
  • NFR-3 (explicit). The LLM API must be accessed through environment variables only; no API key may be hardcoded.
  • NFR-4 (explicit). If no LLM API key is available, the AI / LLM page must show a clear message and the rest of the website must remain fully functional.
  • NFR-5 (explicit). The website must be simple to understand, clean and professional, and beginner-friendly for a BCA student; the Deep Learning section must be kept simple because it is for a BCA academic project.
  • NFR-6 (explicit). The website must be responsive, holding at 375px, 768px and 1280px without clipping headlines, numerals, table cells, chart labels or controls.
  • NFR-7 (explicit). All pages must be connected through navigation.
  • NFR-8 (explicit). The project must run with pip install -r requirements.txt followed by streamlit run app.py, and must ship app.py, requirements.txt, README.md and data/career_dataset.csv.
  • NFR-9 (explicit). The final project must have one clear purpose — use the complete Data Science/AI workflow to build a Career Prediction System — while each page demonstrates one specific topic.
  • NFR-10 (required_inference). Dependent modules must not fabricate data: when no usable dataset is selected, a module shows a clear prerequisite message and a route to Data Collection rather than inventing statistics, charts or predictions.
  • NFR-11 (direction). Motion is restrained and functional (160ms ease-out; charts 400ms), with a full prefers-reduced-motion end-state fallback; no gradients, no dark mode, no neon glows or particle fields.
Page 19 of 21

10. Tech Stack

Source-specified technology is preserved exactly:

  • Python — application language.
  • Streamlit — the website framework; the app runs with streamlit run app.py.
  • Pandas — dataset loading, cleaning, transformation and preview.
  • NumPy — numerical operations.
  • Scikit-learn — Random Forest and Logistic Regression models, evaluation metrics, confusion matrix, TF-IDF vectorization, and the MLPClassifier neural network.
  • Matplotlib / Plotly — static and interactive charts on the Data Visualization page.
  • NLP with TF-IDF — text cleaning, tokenization and keyword detection on the NLP page.
  • MLP Neural Network — the Deep Learning / Neural Network page (Scikit-learn MLPClassifier, or TensorFlow/Keras if appropriate).
  • LLM API — the AI Career Assistant, accessed through environment variables only.

Project files: app.py, requirements.txt, README.md, data/career_dataset.csv.

Page 20 of 21

11. Assumptions and Constraints

Constraints (explicit, binding).

  1. Do NOT hardcode any API key; the LLM API must be accessed through environment variables.
  2. If no LLM API key is available, show a clear message and keep the rest of the website fully functional.
  3. No fake buttons.
  4. No placeholder pages.
  5. No hardcoded prediction results.
  6. Real dataset processing and real Machine Learning prediction are required.
  7. All pages must be connected through navigation.
  8. Keep the project simple and understandable for a BCA student; the Deep Learning section must be kept simple because it is for a BCA academic project.
  9. The website must be fully working, simple to understand, clean and professional, beginner-friendly, and responsive.
  10. The final project must have one clear purpose: use the complete Data Science/AI workflow to build a Career Prediction System, while each page demonstrates one specific topic.
  11. Required files: app.py, requirements.txt, README.md, data/career_dataset.csv; the website should run using pip install -r requirements.txt and streamlit run app.py.

Assumptions (narrow, labeled).

  • A-1 (required_inference). The bundled data/career_dataset.csv is the default dataset for every module until the user uploads a replacement on Data Collection; the eight specified columns are present.
  • A-2 (required_inference). A submitted or sample dataset must be processed before dependent statistics, charts, preprocessing, feature engineering, model evaluation, neural-network and prediction views can operate; those views show a prerequisite message otherwise.
  • A-3 (required_inference). The LLM API key is supplied through an environment variable; without it the AI Career Assistant shows its unavailable-key message and the other modules remain functional.
  • A-4 (required_inference). The model trained on the Machine Learning page is the model used for the final career prediction, so the prediction reflects the user's selected model.
  • A-5 (required_inference). The website is a single-session, anonymous application: no account, login or stored user profile is required by any accepted journey.
  • A-6 (direction). The visual system is light-mode only, on the warm paper ground, with the specified Fira Sans typography and signal palette; no dark mode is provided.
  • A-7 (default — not specified by user). The exact number of records and features in the bundled sample dataset, and the exact model accuracy figure, are determined by the shipped dataset and the trained model; the dashboard reports the real computed values rather than fixed numbers.
Page 21 of 21

12. Glossary

  • Pipeline — the eleven-stage Data Science/AI workflow demonstrated across the website: Data Collection → Data Cleaning → EDA → Visualization → Preprocessing → Feature Engineering → Machine Learning → NLP → Deep Learning → AI/LLM → Prediction.
  • Career dataset — the sample dataset at data/career_dataset.csv containing Education, Specialization, Skills, Programming Level, Experience, Projects, Certifications and Career/Job Role.
  • EDA — Exploratory Data Analysis; the descriptive statistics and most-common-value insights shown on the EDA page.
  • One Hot Encoding — the categorical encoding applied to Education during preprocessing.
  • Multi-label Encoding — the encoding applied to the multi-valued Skills field during preprocessing.
  • TF-IDF — the term-frequency/inverse-document-frequency vectorization used on the NLP page after text cleaning and tokenization.
  • Detected Skills — the skills/career keywords identified from the user's free text on the NLP page.
  • Random Forest / Logistic Regression — the two selectable Machine Learning models trained on the Machine Learning page.
  • MLPClassifier — the Scikit-learn multilayer perceptron used for the simple neural network on the Deep Learning / Neural Network page.
  • Confusion Matrix — the evaluation display on the Machine Learning page.
  • Confidence — the percentage shown on the Prediction Result page for the predicted career.
  • Top 3 roles — the three highest-probability career roles shown with their probabilities on the Prediction Result page.
  • AI Career Assistant — the generative assistant on the AI / LLM page, backed by an LLM API accessed through environment variables.
  • Project Report — the academic report page containing the twenty listed sections and a Download Report button.

No completed page designs yet.

Completed design pages will appear here when they are ready to preview.

Home / Dashboard: View project summary
Career Prediction: Set education and specialization
Career Prediction: Select skills and experience
Career Prediction: Set projects and certification
Career Prediction: 1. Press PREDICT CAREER
Career Prediction: 2. Complete missing field
Prediction Result: 3. Read predicted career and confidence
Prediction Result: 4. Read recommended skills and explanation
Prediction Result: 5. Compare top 3 career roles
Prediction Result: 6. Read prediction failure message
Career Prediction: 7. Adjust inputs and predict again
NLP: 1. Enter skills in free text
NLP: View Detected Skills
NLP: 2. Resubmit unprocessable text
Career Prediction: Carry detected skills into form
AI / LLM: Ask a career question
AI / LLM: 1. Read generated answer
AI / LLM: Read unavailable-key message
AI / LLM: 2. Ask another question
Machine Learning: View trained model metrics
Prediction Result: Check prediction reflects model

No completed page designs yet.

Completed design pages will appear here when they are ready to preview.

Home / Dashboard: View project summary
Career Prediction: Set education and specialization
Career Prediction: Select skills and experience
Career Prediction: Set projects and certification
Career Prediction: 1. Press PREDICT CAREER
Career Prediction: 2. Complete missing field
Prediction Result: 3. Read predicted career and confidence
Prediction Result: 4. Read recommended skills and explanation
Prediction Result: 5. Compare top 3 career roles
Prediction Result: 6. Read prediction failure message
Career Prediction: 7. Adjust inputs and predict again
NLP: 1. Enter skills in free text
NLP: View Detected Skills
NLP: 2. Resubmit unprocessable text
Career Prediction: Carry detected skills into form
AI / LLM: Ask a career question
AI / LLM: 1. Read generated answer
AI / LLM: Read unavailable-key message
AI / LLM: 2. Ask another question
Machine Learning: View trained model metrics
Prediction Result: Check prediction reflects model