winter-career

byregi

Create a simple, clean, fully working **AI & Data Science Predictive Website** for my BCA college project. ## Project Title **AI-Based Career Prediction System** The main purpose of the website is to demonstrate the complete Data Science/AI workflow: **Data Collection → Data Cleaning → EDA → Visualization → Preprocessing → Feature Engineering → Machine Learning → NLP → Deep Learning → AI/LLM → Prediction** The final prediction should use the user's **Education, Skills and Experience** to predict a suitable career/job role. --- # WEBSITE PAGES ## 1. Home / Dashboard Show: * Project title * Short project description * Number of records * Number of features * Machine Learning model used * Model accuracy * Navigation to all modules Also show a simple workflow: **Data → Cleaning → Analysis → ML → AI → Prediction** --- ## 2. Data Collection Purpose: Demonstrate how data is collected. Features: * Upload CSV file * View uploaded dataset * Show number of rows and columns * Show column names * Show first 5/10 records * Allow downloading the dataset Use a sample career dataset containing: * Education * Specialization * Skills * Programming Level * Experience * Projects * Certifications * Career/Job Role --- ## 3. Data Cleaning Purpose: Demonstrate how raw/messy data is cleaned. Show: * Missing values * Duplicate records * Incorrect data types * Missing value handling * Duplicate removal * Cleaned dataset Display before/after statistics. Example: **Before Cleaning** * Rows: 500 * Missing values: 25 * Duplicates: 8 **After Cleaning** * Rows: 492 * Missing values: 0 * Duplicates: 0 --- ## 4. Exploratory Data Analysis (EDA) Purpose: Understand the dataset. Show: * Dataset statistics * Mean * Median * Minimum * Maximum * Standard deviation * Most common education * Most common skill * Most common career Also show useful insights from the dataset. Example: > Python is one of the most common skills among Data Analyst records. --- ## 5. Data Visualization Purpose: Represent data using graphs. Create interactive/simple charts such as: * Education distribution * Skills distribution * Career distribution * Experience distribution * Certification distribution * Education vs Career * Skills vs Career Use: * Matplotlib * Plotly Graphs should update based on the selected dataset. --- ## 6. Data Preprocessing Purpose: Prepare data for Machine Learning. Show: * Categorical encoding * Numerical feature processing * Missing value handling * Feature scaling where required * Train/Test split Explain briefly what preprocessing is doing. Example: **Education → One Hot Encoding** **Experience → Numerical Encoding** **Skills → Multi-label Encoding** --- ## 7. Feature Engineering Purpose: Create useful features from existing data. Create features such as: * Number of Skills * Number of Projects * Experience Score * Certification Score * Programming Skill Score Show the new features in a table. Example: **Python + SQL + ML = Skill Count 3** --- ## 8. Machine Learning Purpose: Train a real Machine Learning model. Use: * Random Forest * Logistic Regression Allow the user to select the model. Show: * Training dataset * Testing dataset * Accuracy * Precision * Recall * F1 Score * Confusion Matrix Display the trained model information. The ML model should actually be used for the final career prediction. --- ## 9. NLP Purpose: Demonstrate Natural Language Processing. Add a small text input: **"Enter your skills or career interest:"** Example: > I know Python, SQL and data visualization. Process the text using: * Text cleaning * Tokenization * TF-IDF Then identify relevant skills/career keywords. Show: **Detected Skills:** * Python * SQL * Data Visualization Also provide a simple sentiment analysis demonstration using sample feedback data. --- ## 10. Deep Learning / Neural Network Purpose: Demonstrate a basic neural network. Create a simple neural network using: * Scikit-learn MLPClassifier or TensorFlow/Keras if appropriate. Use the processed career dataset. Show: * Neural network architecture * Training accuracy * Testing accuracy * Loss/accuracy graph if available Keep this section simple because it is for a BCA academic project. --- ## 11. AI / LLM Purpose: Demonstrate Generative AI. Create a simple **AI Career Assistant**. The user can ask questions such as: > Which skills should I learn for Data Science? > What career can I choose after BCA? > How can I improve my Python skills? Use an LLM API through environment variables. If no API key is available, show a clear message and keep the rest of the website fully functional. Do NOT hardcode any API key. --- # 12. Career Prediction — MAIN PAGE This is the most important page. Create a simple form: ### Education * 10th * 12th * Diploma * BCA * B.Tech * MCA * Other ### Specialization * Computer Science * IT * Data Science * AI/ML * Software Engineering * Other ### Skills Allow multiple selections: * Python * Java * C/C++ * JavaScript * HTML/CSS * SQL * Machine Learning * Data Analysis * Data Visualization * AI * Communication * Problem Solving ### Experience * Fresher * <1 Year * 1–2 Years * 2–5 Years * 5+ Years ### Projects * 0 * 1 * 2–3 * 4+ ### Certification * Yes * No Then add a large button: **🔮 PREDICT CAREER** --- # 13. Prediction Result After prediction, show: ### Predicted Career Example: **Data Analyst** ### Confidence **85%** ### Recommended Skills * Python * SQL * Data Analysis * Power BI ### Why this prediction? Show simple explanations based on the user's input. Example: > Your prediction is influenced by your Python, SQL and Data Analysis skills and your BCA background. Also show the **top 3 possible career roles** with their prediction probabilities, without making the interface complicated. --- # 14. Project Report Create a page containing: * Introduction * Problem Statement * Objectives * Dataset * Data Collection * Data Cleaning * EDA * Visualization * Preprocessing * Feature Engineering * Machine Learning * NLP * Deep Learning * AI/LLM * Prediction * Results * Limitations * Future Scope * Conclusion Add a **Download Report** button. --- # TECHNOLOGY Use: * Python * Streamlit * Pandas * NumPy * Scikit-learn * Matplotlib / Plotly * NLP with TF-IDF * MLP Neural Network * LLM API Keep the project simple and understandable for a BCA student. --- # IMPORTANT REQUIREMENTS The website must be: * Fully working * Simple to understand * Clean and professional * Beginner-friendly * Responsive * No fake buttons * No placeholder pages * No hardcoded prediction results * Real dataset processing * Real Machine Learning prediction * All pages connected through navigation Create these files: ```text app.py requirements.txt README.md data/career_dataset.csv ``` The website should run using: ```bash pip install -r requirements.txt streamlit run app.py ``` Make sure the final project has **one clear purpose: use the complete Data Science/AI workflow to build a Career Prediction System**, while each page demonstrates one specific topic.

No preview

Comments (0)

No comments yet. Be the first!

System Requirements

Page 1 of 25

System Requirements Document for winter-career

1. Introduction

winter-career is the working project name for the AI-Based Career Prediction System, a simple, clean, fully working AI & Data Science predictive website built as a BCA college capstone project. Its single purpose is to demonstrate the complete Data Science / AI workflow end to end — Data Collection → Data Cleaning → EDA → Visualization → Preprocessing → Feature Engineering → Machine Learning → NLP → Deep Learning → AI/LLM → Prediction — and to use that workflow to predict a suitable career/job role from a user's Education, Skills and Experience.

The website is a teaching artefact, not a commercial product. It is intended for a BCA student who builds and presents it, a project guide, and an examiner. Every page demonstrates exactly one topic in the pipeline, and the pages are connected through navigation so the whole route reads as one continuous workflow. The interface must be simple to understand, clean and professional, beginner-friendly, responsive, and free of fake buttons, placeholder pages, and hardcoded prediction results. All dataset processing and all Machine Learning prediction must be real.

The project is delivered as a Streamlit application with the files app.py, requirements.txt, README.md, and data/career_dataset.csv, and runs with:

pip install -r requirements.txt
streamlit run app.py
Page 2 of 25

2. System Overview

The system is a single Streamlit application composed of fourteen connected pages. A persistent left module rail lists the eleven workflow stages as numbered entries (01 Data Collection through 11 Prediction), each with its coded colour dot, so the pipeline is always visible and the current stage is always marked. The Home / Dashboard is the public entry surface and introduces the project, its dataset size, its feature count, the Machine Learning model in use, the model accuracy, and the workflow strip.

The application loads or accepts the sample career dataset, then carries that dataset through cleaning, exploratory analysis, visualization, preprocessing, feature engineering, model training, NLP, and a neural network demonstration. The trained Machine Learning model is the model actually used by the Career Prediction form, and the Prediction Result page presents the predicted role, confidence, recommended skills, an explanation grounded in the user's own input, and the top 3 alternative roles with probabilities.

Two accepted human personas use the system: the BCA Student / Project Presenter, who walks the pipeline stage by stage and presents the report, and the Career Seeker / Prediction User, who submits the prediction form, describes skills in the NLP text input, and asks the AI Career Assistant questions. All fourteen pages are reachable without an account; the application owns no login, no user profile, and no stored personal history.

The AI / LLM page uses an LLM API through environment variables. No API key is ever hardcoded. If no API key is available, the page shows a clear message and the rest of the website remains fully functional.

Narrow exclusions. The project does not include user accounts, authentication, role-based permissions, payment, or any commercial career-placement service. It does not include stock photography, 3D renders, or decorative illustration. It does not include any capability beyond the eleven workflow stages, the prediction form and result, and the project report.

Page 3 of 25

2a. Product Interpretation and Delivery Boundary

Everything the user sees is first-party Streamlit UI owned by this application. There is no provider-owned surface, no external destination, and no headless delivery: the entire product is the running Streamlit app, and every page is reachable from the navigation rail and the Home dashboard.

The current delivery horizon covers all fourteen pages and the full pipeline. The only external dependency is the optional LLM API used by the AI Career Assistant; that dependency is read from environment variables at runtime, and its absence degrades exactly one page's assistant responses while leaving every other page and every other capability working. The dataset is a local CSV file shipped with the project and can also be replaced by a user-uploaded CSV on the Data Collection page.

Nothing in this document is deferred to a future release. The Project Report page's "Future Scope" section is report content describing possible extensions, not a commitment to build them.

2b. Source Content Inventory

No reference directive with content_source authority was supplied, so no source content inventory is included. The sample career dataset is defined by the user's own field list and is specified in the Functional Requirements and Page Content coverage below.

2c. Page Content and Component Coverage

Page 4 of 25

Home / Dashboard

  • Information and state: project title "AI-Based Career Prediction System"; a short project description; number of records; number of features; the Machine Learning model used; the model accuracy; the workflow strip Data → Cleaning → Analysis → ML → AI → Prediction; navigation to all modules.
  • Primary actions: navigate to any module from the workflow strip or the module rail; open the Career Prediction page from the burnt-orange PREDICT CAREER control.
  • Supporting actions: read the stat band; read the workflow strip; follow the rail to the current stage.
  • Domain entities: dataset summary (record count, feature count), model summary (model name, accuracy).
  • Component responsibilities: typographic masthead with kicker, headline, and one-paragraph purpose statement; ruled stat band of four label/value pairs in Fira Mono with tabular numerals; eleven-node workflow strip with numbered labels, coded colour dots, and a 2px hairline connector; persistent left module rail with numbered entries and coded dots.
  • States: loading — stat band and workflow strip show placeholder rules while the dataset and model summary are computed; empty — if the dataset cannot be loaded, the stat band shows an explicit unavailable message and the workflow strip still renders so navigation remains usable; success — all four stats populated and the workflow strip's active connector drawn; error — a clear message naming the failed step (dataset load or model summary) with a retry action; recovery — retry reloads the dataset and recomputes the summary without leaving the page.

Data Collection

  • Information and state: the uploaded or sample dataset rendered as a ruled table; number of rows and columns; column names; the first 5/10 records; the download control for the dataset.
  • Primary actions: upload a CSV file; download the dataset.
  • Supporting actions: switch the preview between the first 5 and first 10 records; scroll the column-name list.
  • Domain entities: career dataset with the columns Education, Specialization, Skills, Programming Level, Experience, Projects, Certifications, Career/Job Role.
  • Component responsibilities: file uploader; dataset preview table with ruled rows and tabular numerals; row/column count readout; column-name list; record-preview selector; download control.
  • States: loading — uploader and preview show a working indicator while the CSV is parsed; empty — before any upload, the sample dataset from data/career_dataset.csv is shown so the page is never blank; success — the dataset renders with its counts, column names, and preview rows; error — a malformed or unreadable CSV produces a clear message naming the parse failure and the sample dataset remains available; recovery — the user can re-upload a corrected file or continue with the sample dataset.

Data Cleaning

  • Information and state: missing values; duplicate records; incorrect data types; missing value handling; duplicate removal; the cleaned dataset; before/after statistics in the form Before Cleaning — Rows, Missing values, Duplicates and After Cleaning — Rows, Missing values, Duplicates.
  • Primary actions: run the cleaning operation; inspect the cleaned dataset.
  • Supporting actions: compare the before and after statistic blocks; inspect which columns carried missing values or duplicates.
  • Domain entities: raw dataset, cleaned dataset, cleaning statistics.
  • Component responsibilities: before/after statistic panels; missing-value report; duplicate report; data-type report; cleaned-dataset table.
  • States: loading — statistic panels show a working indicator while cleaning runs; empty — before cleaning runs, the panels show the raw dataset's counts with the after block marked as not yet computed; success — both statistic blocks populated and the cleaned dataset rendered; error — a failure in cleaning reports which operation failed and leaves the raw dataset intact; recovery — the user can re-run cleaning after correcting the input on Data Collection.
Page 5 of 25

Exploratory Data Analysis (EDA)

  • Information and state: dataset statistics — mean, median, minimum, maximum, standard deviation; most common education; most common skill; most common career; useful insights from the dataset, for example "Python is one of the most common skills among Data Analyst records."
  • Primary actions: review the statistics and insights.
  • Supporting actions: read the most-common values as separate labelled rows.
  • Domain entities: cleaned dataset, descriptive statistics, most-common categorical values, generated insights.
  • Component responsibilities: statistics table with ruled label/value rows; most-common-value panel; insight list.
  • States: loading — statistics table shows a working indicator while the cleaned dataset is summarized; empty — if the cleaned dataset is unavailable, the page states that cleaning must run first and links to Data Cleaning; success — statistics, most-common values, and insights all populated; error — a computation failure names the statistic that failed and preserves the values already computed; recovery — re-running the summary after cleaning succeeds restores the full page.

Data Visualization

  • Information and state: charts for Education distribution, Skills distribution, Career distribution, Experience distribution, Certification distribution, Education vs Career, and Skills vs Career.
  • Primary actions: select a chart to view; interact with Plotly charts.
  • Supporting actions: switch between Matplotlib and Plotly renderings where both are provided; change the selected dataset so the charts update.
  • Domain entities: cleaned dataset, chart series.
  • Component responsibilities: chart selector; Matplotlib figure renderer; Plotly interactive chart renderer; chart captions.
  • States: loading — the selected chart shows a working indicator while it is built from the current dataset; empty — if no dataset is available, the page states that a dataset must be loaded and links to Data Collection; success — the selected chart renders against the currently selected dataset; error — a chart that cannot be built reports the failing chart name and the reason, and the other charts remain selectable; recovery — selecting another chart or reloading the dataset rebuilds the failed chart.

Data Preprocessing

  • Information and state: categorical encoding; numerical feature processing; missing value handling; feature scaling where required; train/test split; brief explanations of what preprocessing is doing, for example Education → One Hot Encoding, Experience → Numerical Encoding, Skills → Multi-label Encoding.
  • Primary actions: run preprocessing; inspect the resulting split.
  • Supporting actions: read the encoding explanation for each column.
  • Domain entities: cleaned dataset, encoded feature matrix, scaled numerical features, train split, test split.
  • Component responsibilities: encoding explanation panel; encoded-feature preview; scaling summary; train/test split summary with row counts.
  • States: loading — the split summary shows a working indicator while encoding and splitting run; empty — before preprocessing runs, the page shows the planned encodings with the split marked as not yet computed; success — encodings, scaling, and split row counts all populated; error — a failure names the failing step (encoding, scaling, or split) and leaves the cleaned dataset untouched; recovery — re-running preprocessing after correcting the input restores the page.
Page 6 of 25

Feature Engineering

  • Information and state: the created features — Number of Skills, Number of Projects, Experience Score, Certification Score, Programming Skill Score — shown in a table, with the worked example Python + SQL + ML = Skill Count 3.
  • Primary actions: run feature engineering; inspect the new-feature table.
  • Supporting actions: read the worked example explaining how a feature is counted.
  • Domain entities: cleaned dataset, engineered feature table.
  • Component responsibilities: new-feature table with ruled rows and tabular numerals; worked-example panel; feature definition list.
  • States: loading — the feature table shows a working indicator while features are computed; empty — before engineering runs, the page lists the features that will be created with the table marked as not yet computed; success — the new-feature table renders with all five features; error — a failure names the feature that could not be computed and preserves the features already produced; recovery — re-running feature engineering restores the full table.

Machine Learning

  • Information and state: the training dataset; the testing dataset; accuracy; precision; recall; F1 score; confusion matrix; trained model information; the model selector offering Random Forest and Logistic Regression.
  • Primary actions: select the model (Random Forest or Logistic Regression); train the model.
  • Supporting actions: inspect the training and testing dataset summaries; read the confusion matrix as a ruled grid; read the trained model information.
  • Domain entities: training dataset, testing dataset, trained model, evaluation metrics, confusion matrix.
  • Component responsibilities: model selector; training and testing dataset summaries; metric readouts for accuracy, precision, recall, and F1 score; confusion matrix grid; trained model information panel.
  • States: loading — metric readouts show a working indicator while training runs; empty — before training, the page shows the selected model and the dataset summaries with metrics marked as not yet computed; success — all five metrics and the confusion matrix populated for the selected model, and the trained model is the one used by Career Prediction; error — a training failure names the model and the reason and leaves any previously trained model in place; recovery — selecting the other model or re-running training restores the page.

NLP

  • Information and state: the text input labelled "Enter your skills or career interest:"; the processed text after text cleaning, tokenization, and TF-IDF; the identified relevant skills/career keywords; the Detected Skills list; a simple sentiment analysis demonstration using sample feedback data.
  • Primary actions: enter skills or career interest text; submit it for processing.
  • Supporting actions: read the detected skills list; read the sentiment analysis demonstration.
  • Domain entities: input text, cleaned tokens, TF-IDF representation, detected skills, sample feedback data with sentiment results.
  • Component responsibilities: text input; processing pipeline display covering cleaning, tokenization, and TF-IDF; detected-skills list; sentiment demonstration panel.
  • States: loading — the detected-skills list shows a working indicator while the text is processed; empty — before any text is entered, the page shows the input with the detected-skills list empty and the sentiment demonstration still available; success — detected skills listed, for example Python, SQL, Data Visualization, alongside the sentiment demonstration results; error — text that yields no recognizable skills reports that outcome clearly rather than showing an empty list without explanation; recovery — entering different text re-runs processing.
Page 7 of 25

Deep Learning / Neural Network

  • Information and state: the neural network architecture; training accuracy; testing accuracy; a loss/accuracy graph if available.
  • Primary actions: train the neural network on the processed career dataset.
  • Supporting actions: read the architecture description; read the loss/accuracy graph.
  • Domain entities: processed career dataset, MLPClassifier (or TensorFlow/Keras) model, architecture description, training accuracy, testing accuracy, loss/accuracy history.
  • Component responsibilities: architecture panel; training and testing accuracy readouts; loss/accuracy graph.
  • States: loading — accuracy readouts show a working indicator while the network trains; empty — before training, the page shows the planned architecture with accuracies marked as not yet computed; success — architecture, training accuracy, testing accuracy, and the loss/accuracy graph all populated; error — a training failure names the reason and states clearly that the loss/accuracy graph is unavailable; recovery — re-running training restores the page.

AI / LLM

  • Information and state: the AI Career Assistant question input; assistant responses; the API availability status; a clear message when no API key is available.
  • Primary actions: ask a career question, for example "Which skills should I learn for Data Science?", "What career can I choose after BCA?", or "How can I improve my Python skills?".
  • Supporting actions: read the API availability status.
  • Domain entities: user question, assistant response, API key availability state.
  • Component responsibilities: question input; response panel; API status indicator; no-key message panel.
  • States: loading — the response panel shows a working indicator while the assistant is queried; empty — before any question, the page shows the input with example questions and the current API availability status; success — the assistant's answer renders in the response panel; error — a failed API call reports the failure clearly and leaves the rest of the website fully functional; recovery — when no API key is present, the page shows a clear message explaining that the assistant is unavailable and that all other pages continue to work, and asking again after a key is configured restores the assistant.

Career Prediction

  • Information and state: the prediction form with Education (10th, 12th, Diploma, BCA, B.Tech, MCA, Other); Specialization (Computer Science, IT, Data Science, AI/ML, Software Engineering, Other); Skills as a multiple selection (Python, Java, C/C++, JavaScript, HTML/CSS, SQL, Machine Learning, Data Analysis, Data Visualization, AI, Communication, Problem Solving); Experience (Fresher, <1 Year, 1–2 Years, 2–5 Years, 5+ Years); Projects (0, 1, 2–3, 4+); Certification (Yes, No); and the large \xf0\x9f\x94\xae PREDICT CAREER button.
  • Primary actions: complete the form; press \xf0\x9f\x94\xae PREDICT CAREER.
  • Supporting actions: review the selected values before submitting.
  • Domain entities: the user's Education, Specialization, Skills, Experience, Projects, and Certification inputs.
  • Component responsibilities: six form controls; the single filled PREDICT CAREER button; validation feedback.
  • States: loading — the button shows a working state while the trained model scores the input; empty — the form renders with no selection made and the button available; success — the prediction is produced and the Prediction Result page is shown; error — an incomplete form or an unavailable trained model produces a clear message naming what is missing and how to fix it; recovery — completing the missing input or training a model on the Machine Learning page makes the button succeed.
Page 8 of 25

Prediction Result

  • Information and state: the predicted career; the confidence percentage; recommended skills; the "Why this prediction?" explanation based on the user's input; the top 3 possible career roles with their prediction probabilities.
  • Primary actions: read the predicted career and confidence; read the recommended skills; read the explanation; read the top 3 alternatives.
  • Supporting actions: return to Career Prediction to change inputs and predict again.
  • Domain entities: predicted career, confidence value, recommended skills, explanation text, top-3 role probabilities.
  • Component responsibilities: predicted-career panel; confidence readout in Fira Mono that counts up once to its final value; recommended-skills list; explanation panel; top-3 alternatives list with probabilities.
  • States: loading — the confidence readout shows a working indicator while the result is prepared; empty — if the page is reached without a prediction, it states that a prediction must be made first and links to Career Prediction; success — predicted career, confidence, recommended skills, explanation, and top 3 alternatives all populated; error — a scoring failure reports the reason and preserves the user's form input so it can be resubmitted; recovery — returning to Career Prediction and pressing \xf0\x9f\x94\xae PREDICT CAREER again produces a fresh result.

Project Report

  • Information and state: the report sections Introduction, Problem Statement, Objectives, Dataset, Data Collection, Data Cleaning, EDA, Visualization, Preprocessing, Feature Engineering, Machine Learning, NLP, Deep Learning, AI/LLM, Prediction, Results, Limitations, Future Scope, and Conclusion; the Download Report button.
  • Primary actions: read the report; download the report.
  • Supporting actions: navigate between report sections.
  • Domain entities: report content, downloadable report file.
  • Component responsibilities: report section renderer with prose measure capped at 68ch; section navigation; Download Report control.
  • States: loading — the report body shows a working indicator while content is assembled; empty — the report always renders its full section list, so no empty state is shown; success — all nineteen sections render and the download produces the report file; error — a download failure reports the reason clearly; recovery — retrying the download after the failure restores the action.
Page 9 of 25

3. Functional Requirements

FR-01 — Project identity and purpose (explicit) As a BCA Student / Project Presenter, I should have a simple, clean, fully working AI & Data Science predictive website titled AI-Based Career Prediction System whose single purpose is to demonstrate the complete Data Science / AI workflow and use it to predict a suitable career/job role, so that the project reads as one coherent capstone rather than a collection of unrelated demos. Lifecycle: initiated by the presenter opening the app; observable result is the Home / Dashboard masthead and workflow strip; failure is a dataset or model summary that cannot be computed, reported on the dashboard with a retry; continuation is navigation into any module. Access: none. Owner: Home / Dashboard.

FR-02 — Complete workflow demonstration (explicit) As a BCA Student / Project Presenter, I should be able to walk the pipeline Data Collection → Data Cleaning → EDA → Visualization → Preprocessing → Feature Engineering → Machine Learning → NLP → Deep Learning → AI/LLM → Prediction, with each page demonstrating exactly one topic, so that the examiner can see every stage of the workflow. Lifecycle: initiated from the module rail or the workflow strip; observable result is the selected stage's page; failure is a stage whose prerequisite data is missing, reported on that page with a link to the prerequisite; continuation is the next stage in the rail. Access: none. Owner: the eleven stage pages plus Home / Dashboard.

FR-03 — Prediction inputs (explicit) As a Career Seeker / Prediction User, I should be able to enter my Education, Skills and Experience — together with Specialization, Projects, and Certification — and receive a predicted suitable career/job role, so that the prediction reflects my own background. Lifecycle: initiated on Career Prediction; observable result is the predicted role on Prediction Result; failure is an incomplete form or an unavailable trained model, reported with the specific missing item; continuation is editing the form and predicting again. Access: none. Owner: Career Prediction → Prediction Result.

FR-04 — Home / Dashboard content (explicit) As a BCA Student / Project Presenter, I should see on the Home / Dashboard the project title, a short project description, the number of records, the number of features, the Machine Learning model used, the model accuracy, navigation to all modules, and the workflow Data → Cleaning → Analysis → ML → AI → Prediction, so that the project's scope and current model state are visible at a glance. Lifecycle: initiated by opening the app; observable result is the populated stat band and workflow strip; failure is an unavailable dataset or model summary, reported explicitly; continuation is navigation to any module. Access: none. Owner: Home / Dashboard.

FR-05 — Data Collection capabilities (explicit) As a BCA Student / Project Presenter, I should be able to upload a CSV file, view the uploaded dataset, see the number of rows and columns, see the column names, see the first 5/10 records, and download the dataset, so that the collection stage is demonstrable and reproducible. Lifecycle: initiated by uploading a CSV or accepting the sample dataset; observable result is the rendered dataset with counts, column names, and preview rows; failure is an unreadable CSV, reported with the parse reason while the sample dataset remains available; continuation is proceeding to Data Cleaning. Access: none. Owner: Data Collection.

FR-06 — Sample career dataset (explicit) As a BCA Student / Project Presenter, I should have a sample career dataset containing Education, Specialization, Skills, Programming Level, Experience, Projects, Certifications, and Career/Job Role, so that every downstream stage has real data to process. Lifecycle: initiated by the application loading data/career_dataset.csv; observable result is the dataset available on Data Collection and usable by every later stage; failure is a missing or unreadable file, reported on Data Collection with the option to upload a replacement; continuation is the cleaning stage. Access: none. Owner: Data Collection (with the file shipped in the project).

FR-07 — Data Cleaning capabilities (explicit) As a BCA Student / Project Presenter, I should see missing values, duplicate records, incorrect data types, missing value handling, duplicate removal, and the cleaned dataset, with before/after statistics in the form Before Cleaning — Rows, Missing values, Duplicates and After Cleaning — Rows, Missing values, Duplicates, so that the cleaning stage is visibly evidenced. Lifecycle: initiated by running cleaning; observable result is the two statistic blocks and the cleaned dataset; failure is a cleaning operation that cannot complete, reported by operation name with the raw dataset preserved; continuation is EDA on the cleaned dataset. Access: none. Owner: Data Cleaning.

FR-08 — EDA capabilities (explicit) As a BCA Student / Project Presenter, I should see dataset statistics — mean, median, minimum, maximum, standard deviation — plus the most common education, most common skill, most common career, and useful insights such as "Python is one of the most common skills among Data Analyst records.", so that the dataset is understood before modeling. Lifecycle: initiated by opening EDA after cleaning; observable result is the statistics table, most-common values, and insights; failure is a statistic that cannot be computed, reported by name with the others preserved; continuation is Data Visualization. Access: none. Owner: Exploratory Data Analysis (EDA).

FR-09 — Data Visualization capabilities (explicit) As a BCA Student / Project Presenter, I should see charts for Education distribution, Skills distribution, Career distribution, Experience distribution, Certification distribution, Education vs Career, and Skills vs Career, built with Matplotlib and Plotly, updating based on the selected dataset, so that the distributions and relationships are visible. Lifecycle: initiated by selecting a chart; observable result is the chart rendered against the currently selected dataset; failure is a chart that cannot be built, reported by chart name while the others remain selectable; continuation is Data Preprocessing. Access: none. Owner: Data Visualization.

FR-10 — Data Preprocessing capabilities (explicit) As a BCA Student / Project Presenter, I should see categorical encoding, numerical feature processing, missing value handling, feature scaling where required, and the train/test split, with brief explanations such as Education → One Hot Encoding, Experience → Numerical Encoding, and Skills → Multi-label Encoding, so that the preparation for Machine Learning is explicit. Lifecycle: initiated by running preprocessing; observable result is the encodings, scaling summary, and split row counts; failure is a failing step named explicitly with the cleaned dataset untouched; continuation is Feature Engineering. Access: none. Owner: Data Preprocessing.

FR-11 — Feature Engineering capabilities (explicit) As a BCA Student / Project Presenter, I should have features created — Number of Skills, Number of Projects, Experience Score, Certification Score, Programming Skill Score — shown in a table, with the worked example Python + SQL + ML = Skill Count 3, so that the engineered inputs to the models are inspectable. Lifecycle: initiated by running feature engineering; observable result is the new-feature table; failure is a feature that cannot be computed, reported by name with the others preserved; continuation is Machine Learning. Access: none. Owner: Feature Engineering.

FR-12 — Machine Learning training and evaluation (explicit) As a BCA Student / Project Presenter, I should be able to select Random Forest or Logistic Regression, train a real model, and see the training dataset, testing dataset, accuracy, precision, recall, F1 score, confusion matrix, and trained model information, so that the model's performance is evidenced. Lifecycle: initiated by selecting a model and training; observable result is the five metrics and the confusion matrix for the selected model; failure is a training failure named by model with any previously trained model preserved; continuation is using the trained model for prediction. Access: none. Owner: Machine Learning.

FR-13 — Trained model used for prediction (explicit) As a Career Seeker / Prediction User, I should have my Career Prediction scored by the Machine Learning model actually trained on the Machine Learning page, so that no prediction result is hardcoded. Lifecycle: initiated by pressing \xf0\x9f\x94\xae PREDICT CAREER; observable result is a prediction produced by the trained model; failure is an unavailable trained model, reported with a link to Machine Learning; continuation is training a model and predicting again. Access: none. Owner: Career Prediction, supported by Machine Learning.

FR-14 — NLP text processing (explicit) As a Career Seeker / Prediction User, I should be able to enter my skills or career interest in the text input labelled "Enter your skills or career interest:" — for example "I know Python, SQL and data visualization." — and have it processed using text cleaning, tokenization, and TF-IDF to identify relevant skills/career keywords, so that free text becomes a detected-skills list such as Python, SQL, Data Visualization. Lifecycle: initiated by submitting text; observable result is the Detected Skills list; failure is text yielding no recognizable skills, reported clearly rather than as a silent empty list; continuation is entering different text or moving to the AI Career Assistant. Access: none. Owner: NLP.

FR-15 — Sentiment analysis demonstration (explicit) As a BCA Student / Project Presenter, I should see a simple sentiment analysis demonstration using sample feedback data on the NLP page, so that the NLP stage covers both keyword extraction and sentiment. Lifecycle: initiated by opening the NLP page; observable result is the sentiment results over the sample feedback data; failure is a sentiment computation that cannot complete, reported clearly; continuation is the Deep Learning stage. Access: none. Owner: NLP.

FR-16 — Neural network demonstration (explicit) As a BCA Student / Project Presenter, I should see a simple neural network built with Scikit-learn MLPClassifier (or TensorFlow/Keras if appropriate) using the processed career dataset, showing the neural network architecture, training accuracy, testing accuracy, and a loss/accuracy graph if available, kept simple for a BCA academic project. Lifecycle: initiated by training the network; observable result is the architecture, both accuracies, and the loss/accuracy graph; failure is a training failure reported with the graph explicitly marked unavailable; continuation is the AI / LLM stage. Access: none. Owner: Deep Learning / Neural Network.

FR-17 — AI Career Assistant (explicit) As a Career Seeker / Prediction User, I should be able to ask the AI Career Assistant questions such as "Which skills should I learn for Data Science?", "What career can I choose after BCA?", and "How can I improve my Python skills?", and receive an answer from an LLM API, so that generative AI guidance is demonstrated. Lifecycle: initiated by submitting a question; observable result is the assistant's answer; failure is an API error reported clearly; continuation is asking another question. Access: none. Owner: AI / LLM.

FR-18 — LLM API key handling (explicit) As a BCA Student / Project Presenter, I should have the LLM API key read from environment variables with no key hardcoded anywhere, and when no API key is available I should see a clear message while the rest of the website remains fully functional, so that the project is safe to share and still demonstrable without a key. Lifecycle: initiated by the application checking the environment at runtime; observable result is either a working assistant or the clear no-key message; failure is a missing key, which is exactly the handled case; continuation is every other page continuing to work normally. Access: none. Owner: AI / LLM.

FR-19 — Career Prediction form (explicit) As a Career Seeker / Prediction User, I should complete a form with Education (10th, 12th, Diploma, BCA, B.Tech, MCA, Other), Specialization (Computer Science, IT, Data Science, AI/ML, Software Engineering, Other), Skills as a multiple selection (Python, Java, C/C++, JavaScript, HTML/CSS, SQL, Machine Learning, Data Analysis, Data Visualization, AI, Communication, Problem Solving), Experience (Fresher, <1 Year, 1–2 Years, 2–5 Years, 5+ Years), Projects (0, 1, 2–3, 4+), and Certification (Yes, No), and then press the large \xf0\x9f\x94\xae PREDICT CAREER button, so that my background is captured exactly as the model expects. Lifecycle: initiated by completing the form; observable result is the submitted prediction request; failure is an incomplete form, reported with the missing field named; continuation is the Prediction Result. Access: none. Owner: Career Prediction.

FR-20 — Prediction Result content (explicit) As a Career Seeker / Prediction User, I should see the predicted career, the confidence percentage, recommended skills, a "Why this prediction?" explanation based on my input — for example "Your prediction is influenced by your Python, SQL and Data Analysis skills and your BCA background." — and the top 3 possible career roles with their prediction probabilities, presented without making the interface complicated, so that the result is both actionable and explainable. Lifecycle: initiated by a successful prediction; observable result is the predicted career, confidence, recommended skills, explanation, and top 3 alternatives; failure is a scoring failure reported with the form input preserved; continuation is returning to Career Prediction to predict again. Access: none. Owner: Prediction Result.

FR-21 — Project Report content and download (explicit) As a BCA Student / Project Presenter, I should have a Project Report page containing Introduction, Problem Statement, Objectives, Dataset, Data Collection, Data Cleaning, EDA, Visualization, Preprocessing, Feature Engineering, Machine Learning, NLP, Deep Learning, AI/LLM, Prediction, Results, Limitations, Future Scope, and Conclusion, with a Download Report button, so that the written submission can be read in the app and downloaded. Lifecycle: initiated by opening the report; observable result is all nineteen sections rendered and a working download; failure is a download failure reported clearly; continuation is retrying the download or navigating away. Access: none. Owner: Project Report.

FR-22 — Navigation connectivity (explicit) As a BCA Student / Project Presenter, I should be able to reach every page through navigation, with the eleven workflow stages listed as numbered entries in the module rail and the current stage marked, so that no page is orphaned and the pipeline reads as one route. Lifecycle: initiated from the rail or the workflow strip; observable result is the destination page with the current stage marked; failure is a stage whose prerequisite is missing, reported on that page with a link to the prerequisite; continuation is the next stage. Access: none. Owner: Home / Dashboard and the shared module rail.

FR-23 — Dataset availability before analysis (required_inference) As a BCA Student / Project Presenter, I should have the sample career dataset loaded or provided before any analysis or modeling stage runs, so that the accepted pipeline is executable from the first stage. Lifecycle: initiated by the application loading data/career_dataset.csv or by a CSV upload; observable result is a dataset available to every downstream stage; failure is a missing or unreadable dataset, reported on Data Collection with the upload alternative; continuation is Data Cleaning. Access: none. Owner: Data Collection.

FR-24 — Ordered processing before prediction (required_inference) As a BCA Student / Project Presenter, I should have the selected dataset processed through cleaning, preprocessing, feature engineering, and model training before the final career prediction is available, so that the prediction is genuinely produced by the demonstrated workflow. Lifecycle: initiated by running each stage in order; observable result is a trained model ready for scoring; failure is a stage that has not been run, reported on Career Prediction with a link to the missing stage; continuation is completing that stage and predicting. Access: none. Owner: the pipeline stages, surfaced on Career Prediction.

FR-25 — Trained model reuse (required_inference) As a Career Seeker / Prediction User, I should have the trained machine-learning model reused for my Career Prediction submission and Prediction Result, so that the result reflects the model I can inspect on the Machine Learning page. Lifecycle: initiated by pressing \xf0\x9f\x94\xae PREDICT CAREER; observable result is a prediction attributed to the currently trained model; failure is an unavailable model, reported with a link to Machine Learning; continuation is training and predicting again. Access: none. Owner: Career Prediction → Prediction Result.

FR-26 — Optional LLM key with functional fallback (required_inference) As a BCA Student / Project Presenter, I should have the LLM API key treated as optional and read only from environment variables, with the AI / LLM page remaining functional by clearly reporting unavailable assistance when no key is present, so that the project runs end to end without any secret. Lifecycle: initiated by the runtime environment check; observable result is either assistant answers or the clear unavailable message; failure is the missing key, which is handled rather than fatal; continuation is every other page working normally. Access: none. Owner: AI / LLM.

FR-27 — Runnable project files and dependencies (required_inference) As a BCA Student / Project Presenter, I should have app.py, requirements.txt, README.md, and data/career_dataset.csv present so that the project runs with pip install -r requirements.txt and streamlit run app.py, so that the examiner can reproduce the demonstration. Lifecycle: initiated by installing dependencies and starting the app; observable result is the running application on Home / Dashboard; failure is a missing dependency or file, reported by the install or start command; continuation is the full pipeline walkthrough. Access: none. Owner: the project itself (no page).

Page 10 of 25

4. User Personas

Page 11 of 25

BCA Student / Project Presenter

Product context. This persona is building and demonstrating the AI-Based Career Prediction System as a BCA college project. They are the person who runs the app in front of a guide or examiner, and they need the whole Data Science / AI pipeline to be visible, ordered, and honest — a working instrument rather than a pitch.

Primary goal. Walk the complete workflow — Data Collection, Data Cleaning, EDA, Visualization, Preprocessing, Feature Engineering, Machine Learning, NLP, Deep Learning, AI/LLM, Prediction — and present the Project Report, so that every stage of the capstone is demonstrably real.

Distinct accepted responsibilities. Uploading or accepting the sample career dataset and inspecting its shape, columns, and preview rows; running cleaning and reading the before/after statistics; reading the EDA statistics and most-common values; selecting and reading each visualization; running preprocessing and reading the encoding explanations; running feature engineering and reading the new-feature table; selecting Random Forest or Logistic Regression and reading accuracy, precision, recall, F1 score, and the confusion matrix; entering text on the NLP page and reading the detected skills and sentiment demonstration; training the neural network and reading its architecture, accuracies, and loss/accuracy graph; reading the AI Career Assistant's availability status; and reading and downloading the Project Report.

Relevant inputs or decisions. Which CSV to upload or whether to use the sample; which chart to view; which model to select; which text to submit on the NLP page; whether to train the neural network; whether to download the report.

Interactions with other accepted participants. The presenter prepares the dataset and the trained model that the Career Seeker / Prediction User depends on, and can demonstrate the prediction form and result on the seeker's behalf during a presentation.

Observable success. Every stage renders real output from the real dataset, the metrics and charts are populated, the report downloads, and no page is a placeholder.

Page 12 of 25

Career Seeker / Prediction User

Product context. This persona wants a concrete answer to "which career fits my background?" and arrives at the Career Prediction page with their own education, specialization, skills, experience, projects, and certification in mind. They may also describe their skills in free text and ask the assistant for guidance.

Primary goal. Submit their Education, Specialization, Skills, Experience, Projects, and Certification, press \xf0\x9f\x94\xae PREDICT CAREER, and receive a predicted career role with confidence, recommended skills, an explanation they can understand, and the top 3 alternatives.

Distinct accepted responsibilities. Completing the six-field prediction form with the exact allowed values; pressing \xf0\x9f\x94\xae PREDICT CAREER; reading the predicted career and confidence; reading the recommended skills; reading the "Why this prediction?" explanation; reading the top 3 possible career roles with probabilities; entering skills or career interest text on the NLP page and reading the detected skills; and asking the AI Career Assistant career questions.

Relevant inputs or decisions. Their own education level, specialization, multi-selected skills, experience band, project count, and certification status; the free-text description of their skills; the questions they ask the assistant.

Interactions with other accepted participants. The seeker depends on the BCA Student / Project Presenter having loaded the dataset and trained the model; if no model is trained, the seeker is directed to the Machine Learning page rather than shown a fabricated result.

Observable success. A predicted career appears with a confidence percentage, recommended skills, an explanation that names their own inputs, and three ranked alternatives — all produced by the trained model, never hardcoded.

Page 13 of 25

5. Core User Flows

Flow 1 — Presenter opens the project and reads the dashboard (BCA Student / Project Presenter)

  1. The presenter runs pip install -r requirements.txt and streamlit run app.py, then opens the app in a browser.
  2. The application loads data/career_dataset.csv and computes the dataset summary and the current model summary.
  3. Home / Dashboard renders the kicker "BCA CAPSTONE · AI & DATA SCIENCE", the headline "AI-Based Career Prediction System", and a one-paragraph purpose statement.
  4. The ruled stat band shows four label/value pairs in Fira Mono — RECORDS, FEATURES, MODEL, ACCURACY — separated by 1px vertical rules, with the metric numbers counting up once on first paint and then holding still.
  5. The eleven-node workflow strip renders edge to edge with numbered Fira Mono labels 01–11, coded colour dots, and a 2px hairline connector; the active stage's connector draws left-to-right over 400ms.
  6. The presenter reads the workflow Data → Cleaning → Analysis → ML → AI → Prediction and the navigation to all modules.
  7. Failure and recovery: if the dataset or model summary cannot be computed, the stat band shows an explicit unavailable message and the workflow strip still renders so navigation remains usable; the presenter retries and the summary recomputes.
  8. Continuation: the presenter selects stage 01 in the module rail and moves into the pipeline.

Flow 2 — Presenter collects and inspects the dataset (BCA Student / Project Presenter)

  1. The presenter opens Data Collection from the module rail; the current stage is marked by a 3px burnt-orange left rule and a filled dot.
  2. The sample dataset from data/career_dataset.csv is already displayed, so the page is never blank.
  3. The presenter optionally uploads a CSV file to replace it.
  4. The page shows the number of rows and columns, the column names, and the first 5/10 records in a ruled table with tabular numerals; the presenter switches the preview between 5 and 10 records.
  5. The presenter confirms the dataset contains Education, Specialization, Skills, Programming Level, Experience, Projects, Certifications, and Career/Job Role.
  6. The presenter downloads the dataset using the download control.
  7. Failure and recovery: an unreadable CSV produces a clear message naming the parse failure, and the sample dataset remains available; the presenter re-uploads a corrected file or continues with the sample.
  8. Continuation: the presenter moves to stage 02, Data Cleaning.
Page 14 of 25

Flow 3 — Presenter cleans the data and reads before/after statistics (BCA Student / Project Presenter)

  1. The presenter opens Data Cleaning.
  2. Before running cleaning, the page shows the raw dataset's counts with the after block marked as not yet computed.
  3. The presenter runs cleaning; the page reports missing values, duplicate records, and incorrect data types, then applies missing value handling and duplicate removal.
  4. The page displays Before Cleaning — Rows, Missing values, Duplicates and After Cleaning — Rows, Missing values, Duplicates side by side, in the shape of the source example (for example Rows 500 / Missing values 25 / Duplicates 8 before, and Rows 492 / Missing values 0 / Duplicates 0 after).
  5. The cleaned dataset renders as a ruled table.
  6. Failure and recovery: if a cleaning operation fails, the page names the operation and leaves the raw dataset intact; the presenter corrects the input on Data Collection and re-runs cleaning.
  7. Continuation: the presenter moves to stage 03, EDA.

Flow 4 — Presenter explores the data (BCA Student / Project Presenter)

  1. The presenter opens Exploratory Data Analysis (EDA).
  2. The page shows dataset statistics — mean, median, minimum, maximum, standard deviation — as ruled label/value rows in Fira Mono with tabular numerals.
  3. The page shows the most common education, the most common skill, and the most common career as separate labelled rows.
  4. The page shows useful insights from the dataset, for example "Python is one of the most common skills among Data Analyst records."
  5. Failure and recovery: if the cleaned dataset is unavailable, the page states that cleaning must run first and links to Data Cleaning; if a single statistic fails, it is named and the others remain visible.
  6. Continuation: the presenter moves to stage 04, Data Visualization.

Flow 5 — Presenter views the charts (BCA Student / Project Presenter)

  1. The presenter opens Data Visualization.
  2. The presenter selects a chart from Education distribution, Skills distribution, Career distribution, Experience distribution, Certification distribution, Education vs Career, and Skills vs Career.
  3. The chart renders against the currently selected dataset, styled to the palette — warm paper plot backgrounds, ink axes, coded series colours, no default chart chrome — using Matplotlib for the simple figures and Plotly for the interactive ones.
  4. The presenter changes the selected dataset and the charts update.
  5. Failure and recovery: a chart that cannot be built reports its name and the reason while the other charts remain selectable; selecting another chart or reloading the dataset rebuilds it.
  6. Continuation: the presenter moves to stage 05, Data Preprocessing.
Page 15 of 25

Flow 6 — Presenter prepares the data for Machine Learning (BCA Student / Project Presenter)

  1. The presenter opens Data Preprocessing.
  2. Before running, the page shows the planned encodings with the split marked as not yet computed.
  3. The presenter runs preprocessing; the page shows categorical encoding, numerical feature processing, missing value handling, feature scaling where required, and the train/test split with row counts.
  4. The page explains briefly what preprocessing is doing, including Education → One Hot Encoding, Experience → Numerical Encoding, and Skills → Multi-label Encoding.
  5. Failure and recovery: a failing step (encoding, scaling, or split) is named explicitly and the cleaned dataset is left untouched; the presenter corrects the input and re-runs.
  6. Continuation: the presenter moves to stage 06, Feature Engineering.

Flow 7 — Presenter engineers and inspects features (BCA Student / Project Presenter)

  1. The presenter opens Feature Engineering.
  2. Before running, the page lists the features that will be created with the table marked as not yet computed.
  3. The presenter runs feature engineering; the page creates Number of Skills, Number of Projects, Experience Score, Certification Score, and Programming Skill Score.
  4. The new features render in a table, with the worked example Python + SQL + ML = Skill Count 3 shown alongside.
  5. Failure and recovery: a feature that cannot be computed is named and the features already produced are preserved; re-running restores the full table.
  6. Continuation: the presenter moves to stage 07, Machine Learning.

Flow 8 — Presenter trains and evaluates a model (BCA Student / Project Presenter)

  1. The presenter opens Machine Learning.
  2. The presenter selects Random Forest or Logistic Regression from the model selector; the selected model is marked in burnt orange, like a transit line currently being ridden.
  3. The presenter trains the model; the page shows the training dataset and the testing dataset.
  4. The page shows accuracy, precision, recall, and F1 score as ruled label/value rows in Fira Mono with tabular numerals, and the confusion matrix as a ruled grid.
  5. The page displays the trained model information.
  6. Failure and recovery: a training failure names the model and the reason and leaves any previously trained model in place; the presenter selects the other model or re-runs training.
  7. Continuation: the trained model is now the model used by Career Prediction; the presenter moves to stage 08, NLP.
Page 16 of 25

Flow 9 — Presenter and seeker use NLP (BCA Student / Project Presenter; Career Seeker / Prediction User)

  1. The presenter or seeker opens NLP.
  2. They enter text in the input labelled "Enter your skills or career interest:", for example "I know Python, SQL and data visualization."
  3. The page processes the text using text cleaning, tokenization, and TF-IDF, and identifies relevant skills/career keywords.
  4. The page shows Detected Skills: Python, SQL, Data Visualization.
  5. The page also shows a simple sentiment analysis demonstration using sample feedback data.
  6. Failure and recovery: text that yields no recognizable skills is reported clearly rather than as a silent empty list; entering different text re-runs processing.
  7. Continuation: the presenter moves to stage 09, Deep Learning / Neural Network.

Flow 10 — Presenter trains the neural network (BCA Student / Project Presenter)

  1. The presenter opens Deep Learning / Neural Network.
  2. Before training, the page shows the planned architecture with accuracies marked as not yet computed.
  3. The presenter trains a simple neural network using Scikit-learn MLPClassifier (or TensorFlow/Keras if appropriate) on the processed career dataset.
  4. The page shows the neural network architecture, the training accuracy, the testing accuracy, and the loss/accuracy graph if available.
  5. Failure and recovery: a training failure is reported with the reason and the loss/accuracy graph is explicitly marked unavailable; re-running training restores the page.
  6. Continuation: the presenter moves to stage 10, AI / LLM.

Flow 11 — Seeker asks the AI Career Assistant (Career Seeker / Prediction User)

  1. The seeker opens AI / LLM.
  2. The page shows the question input, example questions, and the current API availability status.
  3. The seeker asks a question such as "Which skills should I learn for Data Science?", "What career can I choose after BCA?", or "How can I improve my Python skills?".
  4. The application reads the LLM API key from environment variables and sends the question; the assistant's answer renders in the response panel.
  5. Failure and recovery: if no API key is available, the page shows a clear message explaining that the assistant is unavailable and that the rest of the website remains fully functional; every other page continues to work. If an API call fails, the failure is reported clearly and the seeker can ask again.
  6. Continuation: the seeker moves to stage 11, Career Prediction.
Page 17 of 25

Flow 12 — Seeker submits the prediction form (Career Seeker / Prediction User)

  1. The seeker opens Career Prediction, the main page.
  2. The seeker selects Education from 10th, 12th, Diploma, BCA, B.Tech, MCA, Other.
  3. The seeker selects Specialization from Computer Science, IT, Data Science, AI/ML, Software Engineering, Other.
  4. The seeker multi-selects Skills from Python, Java, C/C++, JavaScript, HTML/CSS, SQL, Machine Learning, Data Analysis, Data Visualization, AI, Communication, Problem Solving.
  5. The seeker selects Experience from Fresher, <1 Year, 1–2 Years, 2–5 Years, 5+ Years.
  6. The seeker selects Projects from 0, 1, 2–3, 4+.
  7. The seeker selects Certification from Yes or No.
  8. The seeker presses the large \xf0\x9f\x94\xae PREDICT CAREER button — the single filled element in the interface, a 6px-radius burnt-orange block set flush-left under the form.
  9. Failure and recovery: an incomplete form names the missing field; an unavailable trained model is reported with a link to Machine Learning so the seeker or presenter can train one; the form input is preserved so nothing must be re-entered.
  10. Continuation: the Prediction Result page is shown.

Flow 13 — Seeker reads the prediction result (Career Seeker / Prediction User)

  1. The seeker arrives at Prediction Result after a successful prediction.
  2. The page shows the Predicted Career, for example Data Analyst.
  3. The page shows the Confidence, for example 85%, in a Fira Mono readout that counts up once to its final value and then holds still.
  4. The page shows Recommended Skills, for example Python, SQL, Data Analysis, Power BI.
  5. The page shows Why this prediction? with a simple explanation based on the seeker's input, for example "Your prediction is influenced by your Python, SQL and Data Analysis skills and your BCA background."
  6. The page shows the top 3 possible career roles with their prediction probabilities, presented without making the interface complicated.
  7. Failure and recovery: if the page is reached without a prediction, it states that a prediction must be made first and links to Career Prediction; a scoring failure reports the reason and preserves the form input.
  8. Continuation: the seeker returns to Career Prediction to change inputs and predict again, or moves to the Project Report.
Page 18 of 25

Flow 14 — Presenter reads and downloads the Project Report (BCA Student / Project Presenter)

  1. The presenter opens Project Report.
  2. The page renders the report sections in order: Introduction, Problem Statement, Objectives, Dataset, Data Collection, Data Cleaning, EDA, Visualization, Preprocessing, Feature Engineering, Machine Learning, NLP, Deep Learning, AI/LLM, Prediction, Results, Limitations, Future Scope, Conclusion.
  3. Report prose is set with a measure capped at 68ch so it stays readable.
  4. The presenter presses Download Report and receives the report file.
  5. Failure and recovery: a download failure reports the reason clearly and the presenter retries.
  6. Continuation: the presenter has completed the full pipeline walkthrough and the written submission.
Page 19 of 25

6. Visuals Colors and Theme

The visual direction is typographic infrastructure for a career prediction engine — wayfinding clarity, warm ground, signal-colour pipeline — after Erik Spiekermann. The project is a teaching artefact, so it reads like a well-set technical manual: numbered, gridded, honest, and warm enough to sit with for an hour. Typography is treated as infrastructure, and the eleven pipeline stages are treated like transit lines.

Colour tokens (light mode).

RoleHexUse
Background#F4EFE6Warm paper ground; never pure white
Surface#FFFBF3Even warmer card surface
Text#1C1A17Warm near-black ink, ≈14:1 contrast at body size
Primary#C2410CBurnt signal orange: active pipeline stage, PREDICT CAREER button, selected model
Accent#1F6F5CDeep transit green: confirmed/positive states — accuracy figures, "cleaned", "certified"
Muted#8A8175Labels, metadata, rules
Border rule#E2D9C91px panel and table borders
Coded wayfinding — red#C2410C4px rules, node dots, small caps labels only
Coded wayfinding — yellow#D9A4004px rules, node dots, small caps labels only
Coded wayfinding — green#1F6F5C4px rules, node dots, small caps labels only
Coded wayfinding — blue#2A5C8A4px rules, node dots, small caps labels only; never the primary or button colour

Each of the eleven workflow modules carries one of the four coded tag colours, used only as 4px rules, node dots, and small caps labels — never as large fills. Blue appears only as one of the four coded wayfinding accents.

Typography. Headings use Fira Sans at 700 and 800 for display headings with tight tracking (−0.02em) and sentence case for page titles, and 500 for section headings. A second voice, Fira Mono, carries all data, metrics, column names, code, and pipeline labels, set in uppercase with +0.08em tracking so numbers read like a timetable. Body text is Fira Sans. Headings are flush-left and ragged-right, never centred.

Type scale (1.250 modular ratio): 64 / 48 / 34 / 26 / 20 / 17 / 15 / 13. Display page titles use clamp(40px, 7vw, 64px) from mobile to desktop; section headings 26–34px; body 17px with 1.6 leading; mono labels 13px uppercase. Measure is capped at 68ch for report prose and 46ch for captions.

Shape language. Rectilinear and honest. Maximum 2px corner radius on cards and 6px on buttons — no pill shapes, no blobs, no soft-shadow float. Every panel is a bordered rectangle with a 1px rule in #E2D9C9 and a 3px top edge in its stage's coded colour. Rules are drawn, not implied: horizontal hairlines separate every label/value pair, every table row, and every metric block. Pipeline node dots are 10px flat circles with no glow.

Layout. A visible 12-column grid with a persistent left rail — 240px on desktop, collapsing to a horizontal scrollable module strip on mobile — listing the eleven workflow stages as numbered entries, 01 Data Collection through 11 Prediction, each with its coded colour dot and the current stage marked by a burnt-orange 3px left rule and a filled dot. Content sits in a single wide column with a right-hand metadata gutter on desktop for row counts, feature counts, and model stats. The Home page leads with a full-width workflow strip: eleven numbered nodes connected by a 2px hairline, horizontally scrollable on mobile with each node fully readable as it passes. Tables are the primary visual object — ruled, tabular numerals, alternating warm-paper row tints.

Imagery. Diagrammatic and documentary, never decorative. The Matplotlib and Plotly charts are the imagery, styled to the palette: warm paper plot backgrounds, ink axes, coded series colours, no default chart chrome. Supporting visuals are schematic — the pipeline diagram, an encoding table, a confusion matrix rendered as a ruled grid, a TF-IDF token list set in Fira Mono. No stock photography, no 3D renders, no illustration for its own sake; the interface itself is the visual content.

Avoid. Any blue or indigo as the primary or button colour — #0057FF, #2563EB, #4F46E5, #6366F1, #7C3AED and neighbours are banned; blue appears only as one of four small coded wayfinding accents. Pure white grounds and pure black text. Inter, Roboto, Arial, Helvetica, Open Sans, Lato, Poppins, and system-ui for headings or body — Fira Sans and Fira Mono only, with monospace reserved for data and labels. Gradient-blob heroes, glassmorphism, frosted panels, and decorative blur. Grids of identical hover-lift cards with soft drop shadows — panels are bordered, ruled, and flat. Pill-shaped buttons, blob shapes, and corner radii above 6px. Centred hero stacks and decorative looping animation or parallax. Stock photography, 3D renders, and illustration used as decoration rather than as a diagram. The generic indigo/blue-on-white SaaS template is forbidden.

Page 20 of 25

7. Signature Design Concept

The Home hero is a full-width, left-aligned typographic masthead on warm paper (#F4EFE6) — not a centred SaaS stack. A 13px Fira Mono uppercase kicker reads BCA CAPSTONE · AI & DATA SCIENCE above a 64px Fira Sans 800 headline, AI-Based Career Prediction System, that spans the viewport width and wraps to two lines on desktop and three on mobile at clamp(40px, 7vw, 64px). Directly beneath, a single 46ch paragraph of body copy explains the one purpose of the project. Below that sits a horizontal ruled stat band of four label/value pairs in Fira Mono — RECORDS 500 · FEATURES 8 · MODEL RANDOM FOREST · ACCURACY 87.4% — separated by 1px vertical rules, with the numbers counting up once on first paint.

The signature element is the eleven-node workflow strip running edge to edge beneath the stat band: each node is a numbered Fira Mono label with its coded colour dot, connected by a 2px hairline, with the final PREDICTION node set in burnt orange and brighter than the rest. On load, the active stage's connector draws left-to-right over 400ms. On mobile the strip is horizontally scrollable (overflow-x: auto) so every node becomes fully readable as it passes. There is no gradient, no blob, and no blue button; the only filled element on the screen is the burnt-orange PREDICT CAREER control in the top-right of the content column.

The concept recomposes only accepted content and controls — the project title, the description, the four dashboard statistics, the workflow stages, and the navigation into them. It introduces no new behaviour, page, or destination.

Page 21 of 25

8. Interaction Model & Motion Direction

Interaction Model: Static (direction) Motion Tempo: restrained Hero Dimensionality: flat

Motion is functional and purposeful, in the spirit of transit signage: 160–220ms ease-out on state changes, no bounce, no parallax, no decorative loops. The one expressive moment is the workflow strip, where the active stage's node fills with its coded colour and its connector rule draws left-to-right over 400ms when a page loads. Metric numbers count up once on first paint (600ms, ease-out) and then hold still. Hover on a module rail entry moves the burnt-orange rule 4px; nothing lifts or scales.

Landing Hero Motion Brief

  • Focal subject: the eleven-node workflow strip and the ruled stat band beneath the typographic masthead.
  • Input → transformation → outcome thesis: as the Home page loads, the active stage's connector rule draws left-to-right across the hairline and the four metric values count up once to their final figures; the outcome is a dashboard that reads like a departure board — the pipeline route and the current model state both legible in the first frame.
  • Motion vocabulary: 160–220ms ease-out state changes; a 400ms left-to-right connector draw; a single 600ms ease-out count-up on metrics; a 4px rule shift on rail hover. No bounce, no parallax, no looping.
  • Composed first frame: warm paper ground; 13px Fira Mono kicker; 64px Fira Sans 800 headline wrapping to two lines; a 46ch paragraph; the ruled stat band with four label/value pairs separated by 1px vertical rules; the eleven-node workflow strip running edge to edge with numbered labels, coded dots, and the final PREDICTION node in burnt orange; the burnt-orange PREDICT CAREER control flush in the top-right of the content column.
  • Reduced-motion state: with prefers-reduced-motion, the connector rule is drawn immediately at full length, metric values appear at their final figures without counting up, and the workflow strip wraps into rows or sits in a horizontally scrollable row (overflow-x: auto) whose further nodes are reached by scrolling — every node fully readable, nothing cut off.

No user-requested 3D or WebGL hero was specified, and the direction's hero dimensionality is flat, so no Canvas, R3F, or Drei scene is required.

Page 22 of 25

9. Non-Functional Requirements

NFR-01 — Fully working (explicit) Every page must render real output from the real dataset. No page may be a placeholder, and no control may be a fake button. Rationale: the source states the website must be fully working with no fake buttons and no placeholder pages.

NFR-02 — No hardcoded prediction results (explicit) Prediction results must be produced by the trained Machine Learning model from the user's submitted inputs. No predicted career, confidence value, or probability may be hardcoded. Rationale: explicit source constraint.

NFR-03 — Real dataset processing and real ML prediction (explicit) The dataset must be genuinely parsed, cleaned, analyzed, visualized, preprocessed, and used to train models; the prediction must come from that trained model. Rationale: explicit source constraint.

NFR-04 — No hardcoded API key (explicit) The LLM API key must be read from environment variables only. No key may appear in source, configuration committed to the repository, or the README. Rationale: explicit source constraint.

NFR-05 — Functional without an API key (explicit) When no API key is available, the AI / LLM page must show a clear message and the rest of the website must remain fully functional. Rationale: explicit source constraint.

NFR-06 — Simple and beginner-friendly (explicit) The project must stay simple and understandable for a BCA student: readable code, plain labels, and no unnecessary abstraction. Rationale: explicit source constraint.

NFR-07 — Clean and professional (explicit) The interface must be clean and professional, following the typographic and colour direction in sections 6–8. Rationale: explicit source constraint.

NFR-08 — Responsive (explicit) The website must be responsive. At 375px, 768px, and 1280px, headlines, wordmarks, labels, numbers, cards, and controls must stay entirely inside the viewport and their container, wrapping or scaling to fit, and no other element may cover any part of them. The left module rail collapses to a horizontal scrollable module strip on mobile. The workflow strip is horizontally scrollable on mobile with each node fully readable as it passes. Rationale: explicit source constraint plus the direction's readability rule.

NFR-09 — Navigation connectivity (explicit) All pages must be connected through navigation, with the eleven workflow stages listed as numbered entries and the current stage marked. Rationale: explicit source constraint.

NFR-10 — Reproducible run (explicit) The project must run with pip install -r requirements.txt followed by streamlit run app.py, using the files app.py, requirements.txt, README.md, and data/career_dataset.csv. Rationale: explicit source constraint.

NFR-11 — Single clear purpose (explicit) The final project must have one clear purpose — use the complete Data Science/AI workflow to build a Career Prediction System — while each page demonstrates one specific topic. Rationale: explicit source constraint.

NFR-12 — Accessible contrast and legibility (required_inference) Body text at #1C1A17 on #F4EFE6 must retain its high contrast (≈14:1), and mono labels must remain legible at 13px uppercase with +0.08em tracking. Rationale: required to make the direction's stated contrast and label sizing actually hold in the delivered interface.

Page 23 of 25

10. Tech Stack

All choices below are explicit user requirements.

  • Language: Python.
  • Application framework: Streamlit, with the application entry point in app.py.
  • Data handling: Pandas and NumPy.
  • Machine Learning: Scikit-learn — Random Forest and Logistic Regression for the selectable career prediction models.
  • Visualization: Matplotlib and Plotly.
  • NLP: TF-IDF for text vectorization, alongside text cleaning and tokenization.
  • Deep Learning: Scikit-learn MLPClassifier, or TensorFlow/Keras if appropriate.
  • Generative AI: an LLM API accessed through environment variables.
  • Project files: app.py, requirements.txt, README.md, data/career_dataset.csv.
  • Run commands: pip install -r requirements.txt then streamlit run app.py.

No container, orchestration, or deployment tooling is required by the source; the project runs locally as a Streamlit application.

Page 24 of 25

11. Assumptions and Constraints

Constraints (binding).

  1. Do NOT hardcode any API key; use an LLM API through environment variables.
  2. If no API key is available, show a clear message and keep the rest of the website fully functional.
  3. No fake buttons.
  4. No placeholder pages.
  5. No hardcoded prediction results.
  6. Real dataset processing and real Machine Learning prediction are required.
  7. All pages must be connected through navigation.
  8. Keep the project simple and understandable for a BCA student.
  9. The website must be responsive.
  10. The website must be fully working, clean and professional, and beginner-friendly.
  11. The final project must have one clear purpose — use the complete Data Science/AI workflow to build a Career Prediction System — while each page demonstrates one specific topic.
  12. The project must run with pip install -r requirements.txt and streamlit run app.py.

Assumptions.

  1. [Assumption] The sample dataset data/career_dataset.csv is shipped with the project and contains the eight columns Education, Specialization, Skills, Programming Level, Experience, Projects, Certifications, and Career/Job Role. The source specifies the columns but not the row count; the illustrative figures in the source (500 rows before cleaning, 492 after) are examples of the before/after display, not a mandated dataset size.
  2. [Assumption] The dataset is small enough to be processed in memory by Pandas within a Streamlit session, which is consistent with a BCA capstone demonstration.
  3. [Assumption] The LLM API is reached over the network at runtime and its provider is chosen by whoever configures the environment variable; the source names "an LLM API" without naming a provider, so no provider is fixed here.
  4. [Assumption] The application owns no user accounts, no login, and no stored personal history. Every page is reachable without identity, and no prediction input is persisted beyond the session. This follows from the source, which never asks for accounts and describes a single-session demonstration tool.
  5. [Assumption] The Project Report's "Future Scope" section is report content describing possible extensions; it is not a commitment to build those extensions in the current delivery.
  6. [Assumption] Charts and metrics are computed on demand from the currently selected dataset, so switching datasets updates the visualizations as the source requires.

Presentation and technology defaults.

  • [Default — not specified by user] The exact row count of the shipped sample dataset, beyond the eight required columns.
  • [Default — not specified by user] The specific LLM provider and model behind the environment-variable API key.
  • [Default — not specified by user] The exact train/test split ratio, beyond the requirement that a train/test split is shown.
  • [Default — not specified by user] The specific hyperparameters of the Random Forest, Logistic Regression, and MLPClassifier models, beyond the requirement that they are real trained models.
Page 25 of 25

12. Glossary

  • AI-Based Career Prediction System — the product title of the winter-career project; a Streamlit website that demonstrates the full Data Science/AI workflow and predicts a career role.
  • BCA — Bachelor of Computer Applications; the academic context for which this capstone project is built.
  • Career dataset — the sample CSV at data/career_dataset.csv with the columns Education, Specialization, Skills, Programming Level, Experience, Projects, Certifications, and Career/Job Role.
  • Cleaned dataset — the dataset after missing value handling and duplicate removal on the Data Cleaning page.
  • Confidence — the percentage shown on the Prediction Result page for the predicted career.
  • Confusion matrix — the ruled grid on the Machine Learning page showing predicted versus actual class counts.
  • Detected Skills — the list of skills identified from free text on the NLP page after text cleaning, tokenization, and TF-IDF.
  • EDA — Exploratory Data Analysis; the stage that produces mean, median, minimum, maximum, standard deviation, most common education, most common skill, most common career, and dataset insights.
  • Engineered features — Number of Skills, Number of Projects, Experience Score, Certification Score, and Programming Skill Score, created on the Feature Engineering page.
  • Fira Mono — the monospace typeface used for all data, metrics, column names, code, and pipeline labels.
  • Fira Sans — the humanist sans typeface used for headings and body text.
  • LLM API — the generative AI service used by the AI Career Assistant, accessed through environment variables.
  • MLPClassifier — the Scikit-learn multilayer perceptron used for the Deep Learning / Neural Network demonstration.
  • Module rail — the persistent left navigation listing the eleven workflow stages as numbered entries with coded colour dots.
  • Pipeline — the eleven-stage workflow Data Collection → Data Cleaning → EDA → Visualization → Preprocessing → Feature Engineering → Machine Learning → NLP → Deep Learning → AI/LLM → Prediction.
  • Prediction Result — the page showing the predicted career, confidence, recommended skills, explanation, and top 3 alternative roles.
  • PREDICT CAREER — the single filled burnt-orange button on the Career Prediction page that submits the form.
  • TF-IDF — term frequency–inverse document frequency; the text vectorization method used on the NLP page.
  • Top 3 possible career roles — the three highest-probability alternative roles shown with their probabilities on the Prediction Result page.
  • Train/test split — the division of the preprocessed dataset into training and testing portions, shown on the Data Preprocessing and Machine Learning pages.
  • Workflow strip — the full-width eleven-node diagram on the Home hero, with numbered Fira Mono labels, coded colour dots, and a 2px hairline connector.

No completed page designs yet.

Completed design pages will appear here when they are ready to preview.

Home / Dashboard: 1. Read project purpose and stats
Home / Dashboard: Follow workflow strip to a module
Home / Dashboard: 2. Read unavailable dataset message
Home / Dashboard: 3. Retry dataset load
Data Collection: Inspect sample dataset
Data Collection: Switch preview to 10 records
Data Collection: 1. Upload replacement CSV
Data Collection: 2. Read parse failure message
Data Collection: Continue with sample dataset
Data Collection: Download dataset
Data Cleaning: 1. Run cleaning operation
Data Cleaning: 2. Compare before/after statistics
Data Cleaning: 3. Read failing operation message
Data Cleaning: 4. Re-run cleaning
Exploratory Data Analysis EDA : 5. Review dataset statistics
Exploratory Data Analysis EDA : Read most common values
Exploratory Data Analysis EDA : 6. Follow link to Data Cleaning
Data Visualization: Select a chart
Data Visualization: 1. Interact with Plotly chart
Data Visualization: Switch selected dataset
Data Visualization: 2. Select another chart after failure
Data Preprocessing: Run preprocessing
Data Preprocessing: Read encoding explanations
Data Preprocessing: Read failing step message
Data Preprocessing: Re-run preprocessing
Feature Engineering: Run feature engineering
Feature Engineering: 1. Inspect new-feature table
Feature Engineering: Read worked example
Feature Engineering: 2. Re-run feature engineering
Machine Learning: Select Random Forest or Logistic Regression
Machine Learning: 1. Train selected model
Machine Learning: Read metrics and confusion matrix
Machine Learning: 2. Select other model after failure
Machine Learning: Re-run training
NLP: Enter skills or career interest text
NLP: 1. Submit text for processing
NLP: 2. Read no recognizable skills message
NLP: Read detected skills and sentiment
Deep Learning / Neural Network: Train neural network
Deep Learning / Neural Network: Read architecture and accuracies
Deep Learning / Neural Network: Read training failure and missing graph
Deep Learning / Neural Network: Re-run training
AI / LLM: Read API availability status
AI / LLM: Read no-key unavailable message
Career Prediction: 1. Demonstrate prediction form
Career Prediction: 2. Press PREDICT CAREER
Prediction Result: Read predicted career and confidence
Prediction Result: Read explanation and top 3
Prediction Result: 3. Read predict-first message
Project Report: Read all report sections
Project Report: 1. Download report
Project Report: 2. Read download failure and retry

No completed page designs yet.

Completed design pages will appear here when they are ready to preview.

Home / Dashboard: 1. Read project purpose and stats
Home / Dashboard: Follow workflow strip to a module
Home / Dashboard: 2. Read unavailable dataset message
Home / Dashboard: 3. Retry dataset load
Data Collection: Inspect sample dataset
Data Collection: Switch preview to 10 records
Data Collection: 1. Upload replacement CSV
Data Collection: 2. Read parse failure message
Data Collection: Continue with sample dataset
Data Collection: Download dataset
Data Cleaning: 1. Run cleaning operation
Data Cleaning: 2. Compare before/after statistics
Data Cleaning: 3. Read failing operation message
Data Cleaning: 4. Re-run cleaning
Exploratory Data Analysis EDA : 5. Review dataset statistics
Exploratory Data Analysis EDA : Read most common values
Exploratory Data Analysis EDA : 6. Follow link to Data Cleaning
Data Visualization: Select a chart
Data Visualization: 1. Interact with Plotly chart
Data Visualization: Switch selected dataset
Data Visualization: 2. Select another chart after failure
Data Preprocessing: Run preprocessing
Data Preprocessing: Read encoding explanations
Data Preprocessing: Read failing step message
Data Preprocessing: Re-run preprocessing
Feature Engineering: Run feature engineering
Feature Engineering: 1. Inspect new-feature table
Feature Engineering: Read worked example
Feature Engineering: 2. Re-run feature engineering
Machine Learning: Select Random Forest or Logistic Regression
Machine Learning: 1. Train selected model
Machine Learning: Read metrics and confusion matrix
Machine Learning: 2. Select other model after failure
Machine Learning: Re-run training
NLP: Enter skills or career interest text
NLP: 1. Submit text for processing
NLP: 2. Read no recognizable skills message
NLP: Read detected skills and sentiment
Deep Learning / Neural Network: Train neural network
Deep Learning / Neural Network: Read architecture and accuracies
Deep Learning / Neural Network: Read training failure and missing graph
Deep Learning / Neural Network: Re-run training
AI / LLM: Read API availability status
AI / LLM: Read no-key unavailable message
Career Prediction: 1. Demonstrate prediction form
Career Prediction: 2. Press PREDICT CAREER
Prediction Result: Read predicted career and confidence
Prediction Result: Read explanation and top 3
Prediction Result: 3. Read predict-first message
Project Report: Read all report sections
Project Report: 1. Download report
Project Report: 2. Read download failure and retry