retail-jayapura

bySULTHAN FARREL ALGHASSAN official

JAYAPURA RETAIL INTELLIGENCE MASTER PROMPT 1 — FOUNDATION, DATA WORKSPACE & DATA CLEANING Saya ingin membangun aplikasi web bernama: JAYAPURA RETAIL INTELLIGENCE Aplikasi ini nantinya merupakan sistem AI Retail Intelligence untuk membantu bisnis retail elektronik/gadget di Jayapura dan sekitarnya dalam memahami: - penjualan - permintaan pasar - kebutuhan dan keluhan pelanggan - teknologi - inventory - keuangan dasar - laporan Excel - business intelligence - AI business analysis PRINSIP UTAMA: DATA FIRST. DETERMINISTIC ANALYTICS FIRST. AI INTERPRETATION SECOND. Artinya: Python/Pandas menghitung angka dan statistik terlebih dahulu. AI nantinya hanya menginterpretasikan hasil analisis dan membantu menghasilkan insight/rekomendasi. AI tidak boleh mengarang angka. ================================================== MASTER PROMPT INI TERDIRI DARI 3 PHASE ================================================== PHASE 1 — PROJECT FOUNDATION PHASE 2 — DATA WORKSPACE PHASE 3 — DATA CLEANING & VALIDATION Kerjakan ketiga phase tersebut secara BERURUTAN dalam satu pekerjaan. Jangan berhenti setelah Phase 1 atau Phase 2 hanya karena phase tersebut selesai. Tetapi jangan melanjutkan ke modul yang belum termasuk dalam Master Prompt 1. ================================================== PHASE 1 — PROJECT FOUNDATION ================================================== Buat foundation project yang bersih dan profesional. Struktur awal: jayapura-retail-intelligence/ ├── frontend/ ├── backend/ ├── analytics/ ├── ai/ ├── data/ │ ├── raw/ │ ├── processed/ │ └── external/ ├── reports/ ├── tests/ ├── docs/ ├── README.md ├── .gitignore └── .env.example Gunakan arsitektur: Frontend: - React - TypeScript - Vite Backend: - Python - FastAPI Data: - Pandas - NumPy - openpyxl AI: belum diaktifkan pada master prompt ini. Database: belum diperlukan pada tahap ini kecuali benar-benar diperlukan oleh arsitektur. Buat struktur kode yang mudah dikembangkan. Jangan membuat: - login - authentication - AI agent - chatbot - web scraping - real market data - dummy sales data - dummy revenue - dummy customer - dummy inventory - database bisnis - API key - business intelligence palsu ================================================== PHASE 2 — DATA WORKSPACE ================================================== Setelah foundation selesai, bangun DATA WORKSPACE. Tujuan: Pengguna nantinya dapat memasukkan dataset bisnis seperti: CSV XLSX Dataset dapat berupa data penjualan. Contoh field yang nantinya mungkin tersedia: tanggal transaksi_id cabang/POS produk kategori brand unit harga diskon revenue customer_id Tetapi JANGAN membuat data contoh. Sistem harus mampu menangani kondisi: BELUM ADA DATA. Jika belum ada dataset: tampilkan empty state yang jujur. Jangan menampilkan: Rp0 sebagai data bisnis, 100 transaksi, grafik palsu, produk contoh, atau KPI palsu. Jika belum ada data, katakan bahwa dataset belum tersedia. Data Workspace minimal harus mempunyai konsep: 1. Upload CSV 2. Upload XLSX 3. Dataset listing 4. Dataset status 5. Dataset metadata 6. Dataset preview 7. Data validation status Jika fitur upload memerlukan backend API, buat endpoint yang bersih dan aman. Gunakan validasi tipe file. Jangan menerima sembarang file sebagai dataset bisnis. ================================================== PHASE 3 — DATA CLEANING & VALIDATION ================================================== Setelah Data Workspace selesai, buat fondasi data cleaning. Gunakan Python/Pandas. Tujuannya adalah: RAW DATA ↓ VALIDATION ↓ CLEANING ↓ PROCESSED DATA Jangan melakukan analisis bisnis terlebih dahulu. Buat sistem yang dapat mendeteksi masalah seperti: - kolom wajib tidak ada - nama kolom tidak sesuai - tanggal invalid - angka terbaca sebagai text - unit kosong - harga kosong - revenue kosong - duplikasi transaksi - nilai negatif yang mencurigakan - format tanggal berbeda - whitespace - nama kategori tidak konsisten - nama produk tidak konsisten Buat konsep: raw dataset cleaned dataset validation report Sistem harus menjelaskan masalah data kepada pengguna. Contoh: "Kolom tanggal memiliki 12 nilai yang tidak dapat dibaca." Bukan langsung menghapus data secara diam-diam. Untuk setiap transformasi penting, simpan informasi mengenai apa yang dilakukan. ================================================== DATA CONTRACT ================================================== Buat data contract yang jelas. Minimal definisikan: dataset type required columns optional columns data types validation rules Untuk sales dataset, gunakan konsep canonical fields. Contoh: tanggal transaksi_id pos produk kategori brand unit harga diskon revenue customer_id Jangan memaksa semua field harus ada jika memang tidak diperlukan. Bedakan: REQUIRED OPTIONAL DERIVED Contoh: revenue dapat menjadi DERIVED jika unit dan harga tersedia. ================================================== ATURAN ANGKA ================================================== Semua angka yang berasal dari dataset harus benar-benar dihitung dari dataset. Jangan membuat angka untuk demo. Jangan menggunakan random data. Jangan menggunakan angka hardcoded sebagai data bisnis. ================================================== UI ================================================== Buat UI yang sederhana, profesional, dan mudah dikembangkan. Gunakan bahasa Indonesia untuk interface utama. Gunakan format Indonesia: Rupiah: Rp Tanggal: DD/MM/YYYY Timezone: WIB Tetapi jangan menampilkan angka bisnis jika dataset belum tersedia. ================================================== TESTING ================================================== Buat test dasar untuk: - file validation - dataset validation - column validation - data type validation - cleaning pipeline - empty dataset state Pastikan backend Python dapat dijalankan. Pastikan frontend dapat di-build. Pastikan tidak ada error lint yang serius. ================================================== README ================================================== Update README dengan: - tujuan project - arsitektur - struktur folder - data flow - data contract - roadmap - prinsip DATA FIRST - prinsip deterministic analytics - status phase yang sudah selesai ================================================== BATASAN ================================================== JANGAN membangun: Sales Analytics Engine lengkap Market Intelligence Customer Intelligence Technology Intelligence Inventory Analytics Accounting AI Business Analyst Report Generator Final Dashboard Semua itu akan dilakukan pada master prompt berikutnya. ================================================== HASIL AKHIR MASTER PROMPT 1 ================================================== Setelah ketiga phase selesai, berikan: 1. Struktur folder final. 2. Daftar file penting yang dibuat. 3. Penjelasan arsitektur. 4. Penjelasan Data Workspace. 5. Penjelasan cleaning pipeline. 6. Penjelasan data contract. 7. Daftar endpoint yang dibuat. 8. Daftar test. 9. Hasil lint/build/test. 10. Daftar hal yang BELUM dibangun. Jangan membuat data palsu. Jangan mengarang hasil testing. Jika suatu test gagal, laporkan dengan jujur dan perbaiki jika memungkinkan. SETELAH MASTER PROMPT 1 SELESAI, BERHENTI. Jangan mengerjakan Master Prompt 2 sebelum saya memberikan instruksi berikutnya.

No preview

Comments (0)

No comments yet. Be the first!

System Requirements

Page 1 of 28

System Requirements Document for retail-jayapura

1. Introduction

JAYAPURA RETAIL INTELLIGENCE (project name: retail-jayapura) is a web application intended to become an AI Retail Intelligence system for electronics/gadget retail businesses in Jayapura and the surrounding Papua region. Its long-term purpose is to help those businesses understand penjualan, permintaan pasar, kebutuhan dan keluhan pelanggan, teknologi, inventory, keuangan dasar, laporan Excel, business intelligence, and AI business analysis.

This document covers Master Prompt 1 only, which is delivered in three sequential phases within a single body of work:

  • Phase 1 — Project Foundation: a clean, professional project scaffold with a fixed folder structure and a fixed technology stack.
  • Phase 2 — Data Workspace: the ability for a user to bring in their own business datasets (CSV and XLSX), list them, inspect their status, metadata, and preview, and see their validation status.
  • Phase 3 — Data Cleaning & Validation: a Python/Pandas foundation that moves RAW DATA → VALIDATION → CLEANING → PROCESSED DATA, detects a defined set of data problems, explains them to the user, and records what each important transformation did.

The governing principle across all three phases is:

DATA FIRST. DETERMINISTIC ANALYTICS FIRST. AI INTERPRETATION SECOND.

Python/Pandas computes numbers and statistics first. AI will later only interpret analysis results and help produce insight/recommendation. AI must never invent numbers. In this master prompt, AI is not activated at all.

The audience is the retail data analyst who prepares and cleans the shop's own exports, and the retail business owner/manager who needs certainty that any number the system will eventually show comes from a real dataset rather than a demo figure.

Page 2 of 28

2. System Overview

The current delivery is a first-party web workspace with no login and no authentication. It consists of a React + TypeScript + Vite frontend, a Python + FastAPI backend, and a Pandas/NumPy/openpyxl data layer. There is no AI module active, no business database, no API key, and no fabricated business content of any kind.

The accepted behavior in this delivery is:

  • A user opens the workspace and, when no dataset exists, sees an honest empty state stating that no dataset is available — never Rp0 as business data, never "100 transaksi", never a fake chart, never a sample product, never a fake KPI.
  • A user uploads a business dataset as CSV or XLSX. File type is validated; arbitrary files are not accepted as business datasets.
  • A user sees the dataset listing, each dataset's status, its metadata, a preview of its actual rows, and its validation status.
  • A user runs validation and cleaning against a documented data contract, and receives a validation report and a record of important transformations. Data problems are explained to the user in plain sentences; data is never silently deleted.
  • The interface is in Indonesian, uses Rp for Rupiah, DD/MM/YYYY for dates, and WIB as the timezone — but shows no business numbers at all while no dataset is available.

Actors in this delivery are two human personas (Analis Data Retail, Pemilik/Pengelola Bisnis Retail) and the application's own backend/analytics processes. There are no external providers, no outbound recipients, and no third-party surfaces.

Narrow exclusions for this delivery: no login, no authentication, no AI agent, no chatbot, no web scraping, no real market data, no dummy sales data, no dummy revenue, no dummy customer, no dummy inventory, no business database, no API key, no fake business intelligence, and no sample/random/hardcoded business numbers. Also excluded from this delivery and reserved for later master prompts: a complete Sales Analytics Engine, Market Intelligence, Customer Intelligence, Technology Intelligence, Inventory Analytics, Accounting, AI Business Analyst, Report Generator, and Final Dashboard.

Page 3 of 28

2a. Product Interpretation and Delivery Boundary

This delivery is a data custody instrument, not an analytics product. Its whole promise is that the user can trust what the system says about their own files. That promise is why the workspace opens on an honest absence rather than a dashboard, why every count and size is a real measurement of a real file, and why validation findings are stated as sentences about the user's data rather than as decorative alerts.

Delivery ownership. All eleven surfaces in this delivery are first-party application pages. There is no provider-owned surface, no external destination, and no headless-only delivery. The backend and the Pandas cleaning pipeline are system processes that support the human-facing pages; they are never the only owner of a human-facing capability.

Access ownership. Access to the workspace requires no login and no authentication. This is an explicit source constraint, not an omission: the user forbade building login and authentication in this master prompt. Every page in this delivery is therefore reachable without establishing identity, and no page in this delivery owns an identity-establishment interaction. There is no application-owned identity, no session continuity requirement, and no differentiated permission or role-based visibility anywhere in this delivery. The two personas differ by what they need from the workspace, not by what they are permitted to see.

Current vs. future boundary. Current: foundation scaffold, Data Workspace (upload CSV, upload XLSX, listing, status, metadata, preview, validation status), data contract, validation and cleaning pipeline, validation report, transformation record, honest empty states, basic tests, README. Future (explicitly deferred, not built here): the Sales Analytics Engine, Market Intelligence, Customer Intelligence, Technology Intelligence, Inventory Analytics, Accounting, AI Business Analyst, Report Generator, and Final Dashboard — all to be done in later master prompts. AI interpretation is deferred. A business database is deferred unless the architecture genuinely requires one.

Stop condition. After Master Prompt 1 is complete, work stops. Master Prompt 2 is not started before the next instruction is given.

Page 4 of 28

2b. Source Content Inventory

Not applicable. No reference directive in this request declares content_source authority, so no source content inventory is produced. All product facts in this document come from the authoritative user requirement thread and the accepted Planning Scope.

2c. Page Content and Component Coverage

The page inventory below is the closed, ordered page contract for this delivery. Each page appears exactly once.

Landing

  • Purpose and information/state: the anonymous entry surface. It explains what JAYAPURA RETAIL INTELLIGENCE is, what the Data Workspace is for, and — when no dataset exists — states plainly that no dataset is available yet.
  • Primary actions: navigate into the Data Workspace (Unggah Dataset, Daftar Dataset, Validasi & Pembersihan, Kontrak Data).
  • Supporting actions: read the DATA FIRST / deterministic analytics explanation; read the current phase status.
  • Domain entities: none displayed as data. The page must not render any dataset-derived value.
  • Component responsibilities: workspace shell with left rail; display statement of the honest empty state; ruled horizontal rule; one-sentence body copy explaining that CSV or XLSX can be uploaded and that all numbers will be computed from the user's data rather than from examples; primary upload entry; secondary dashed upload target; deliberate empty graphite field on the right.
  • States: empty — the only meaningful state on first use, and the state that must be honest. loading — not applicable, since no dataset is being fetched. success — not applicable on this page. error — not applicable on this page. recovery — not applicable on this page.

Upload CSV

  • Purpose and information/state: the workspace for bringing a CSV business dataset into the system. Shows the accepted file type, the current upload state, and the outcome of file-type validation.
  • Primary actions: select or drop a CSV file; submit it for ingestion.
  • Supporting actions: read the accepted-format note; retry after a rejected file; navigate to the resulting dataset.
  • Domain entities: the uploaded file (name, size, type), the resulting dataset record.
  • Component responsibilities: dashed 1px dropzone that becomes 2px dashed tangerine only while a file is dragged over it; determinate tangerine progress bar during upload (never a spinner); file-type validation feedback; rejection message naming the reason; link to the created dataset on success.
  • States: empty — no file selected yet, dropzone shown plainly. loading — determinate progress bar while the file is transferred and parsed. success — dataset created; the user is offered the dataset's status, metadata, and preview. error — file type rejected, or the file could not be parsed as CSV; the message states which condition occurred. recovery — the user can select a different file and resubmit without losing the page.
Page 5 of 28

Upload XLSX

  • Purpose and information/state: the workspace for bringing an XLSX business dataset into the system. Shows the accepted file type, the current upload state, and the outcome of file-type validation.
  • Primary actions: select or drop an XLSX file; submit it for ingestion.
  • Supporting actions: read the accepted-format note; retry after a rejected file; navigate to the resulting dataset.
  • Domain entities: the uploaded file (name, size, type), the resulting dataset record.
  • Component responsibilities: dashed 1px dropzone that becomes 2px dashed tangerine only while a file is dragged over it; determinate tangerine progress bar during upload; file-type validation feedback; rejection message naming the reason; link to the created dataset on success.
  • States: empty — no file selected yet. loading — determinate progress bar while the file is transferred and parsed. success — dataset created; the user is offered the dataset's status, metadata, and preview. error — file type rejected, or the file could not be parsed as XLSX; the message states which condition occurred. recovery — the user can select a different file and resubmit without losing the page.

Datasets

  • Purpose and information/state: the listing of datasets available in the Data Workspace, with each dataset's identity and current state.
  • Primary actions: open a dataset's status, metadata, preview, and validation status.
  • Supporting actions: start a new upload; filter or scan the listing.
  • Domain entities: dataset record (name, source file, format, upload time, row count, column count, current processing state).
  • Component responsibilities: ruled ledger-style table with tabular numerals right-aligned in ruled columns; row hover as an instant flat surface fill; status chip per row; honest empty state when the listing is empty.
  • States: empty — no dataset has been uploaded; the page states that no dataset is available and offers the upload entry. loading — the listing is being read. success — one or more datasets listed with real measured counts. error — the listing could not be read; the message states the failure. recovery — retry the listing read.

Dataset Status

  • Purpose and information/state: the processing and validation status of a dataset, including where it currently sits in the RAW → VALIDASI → PEMBERSIHAN → PROCESSED flow.
  • Primary actions: run validation; run cleaning; open the validation report; open the data issues list.
  • Supporting actions: return to the dataset listing; open metadata or preview.
  • Domain entities: dataset record, processing state, validation state, cleaning state, timestamps.
  • Component responsibilities: four-stage pipeline strip (RAW → VALIDASI → PEMBERSIHAN → PROCESSED) as four ruled rectangles joined by 1px arrows, with the current stage filled tangerine and the rest hairline-outlined; status chips; action links.
  • States: empty — no dataset selected or no dataset exists; the page states that no dataset is available. loading — status is being read. success — current stage and validation state shown truthfully. error — status could not be read, or a validation/cleaning run failed; the failure is stated. recovery — re-run validation or cleaning from this page.
Page 6 of 28

Dataset Metadata

  • Purpose and information/state: metadata for the selected dataset, including its dataset type and its data structure.
  • Primary actions: inspect the column list and each column's detected type; open the validation report.
  • Supporting actions: return to the dataset listing; open preview.
  • Domain entities: dataset record, column inventory (column name, detected type, role as REQUIRED / OPTIONAL / DERIVED), source file facts.
  • Component responsibilities: sticky left pane in the two-pane dataset detail view; ruled panel header with title left and status chip right; monospace schema listing; DERIVED acid-lime monospace chip beside any canonical field the system computed rather than read.
  • States: empty — no dataset selected; the page states that no dataset is available. loading — metadata is being read. success — metadata and column inventory shown. error — metadata could not be read. recovery — retry the metadata read.

Dataset Preview

  • Purpose and information/state: a preview of the actual rows of the selected dataset. No sample data is ever shown; only rows that exist in the user's file.
  • Primary actions: scroll the preview table; switch between raw and cleaned preview where a cleaned dataset exists.
  • Supporting actions: return to metadata or validation report.
  • Domain entities: dataset rows, column headers, row indices.
  • Component responsibilities: right pane of the two-pane dataset detail view; horizontally scrollable ruled table with a frozen first column; tabular numerals in JetBrains Mono; row hover as an instant flat surface fill.
  • States: empty — no dataset selected, or the selected dataset has no rows; the page states that no dataset is available rather than rendering placeholder rows. loading — preview rows are being read. success — real rows rendered. error — preview could not be read. recovery — retry the preview read.

Validation Status

  • Purpose and information/state: a summary of the dataset's validation state against the data contract rules.
  • Primary actions: open the full validation report; open the data issues list; run validation.
  • Supporting actions: return to the dataset listing.
  • Domain entities: validation result per rule, severity, affected-row counts.
  • Component responsibilities: four-stage pipeline strip with VALIDASI as the current stage; severity chips (tangerine Perlu tindakan, lime Bersih, muted Info); tabular affected-row counts.
  • States: empty — no dataset selected or no dataset exists; the page states that no dataset is available. loading — validation results are being read. success — validation summary shown with real counts. error — validation could not be run or results could not be read. recovery — re-run validation.
Page 7 of 28

Validation Report

  • Purpose and information/state: the report explaining which required columns, optional columns, data types, and validation rules were applied to the dataset, and what the outcome was for each.
  • Primary actions: read each finding; follow the action link on a finding.
  • Supporting actions: open the data issues list; return to validation status.
  • Domain entities: data contract definition (dataset type, required columns, optional columns, data types, validation rules), per-rule outcome, affected-row counts.
  • Component responsibilities: ruled ledger of findings — one row per issue with a severity chip, a monospace affected-row count, a human sentence, and a right-side action link; never a coloured card with a drop shadow; rows stagger in at 40ms intervals on first paint and then never animate again; monospace schema listing of the contract.
  • States: empty — no validation has been run, or no dataset exists; the page states which of the two applies. loading — the report is being produced. success — findings listed with real counts and plain sentences. error — the report could not be produced. recovery — re-run validation.

Data Cleaning

  • Purpose and information/state: the workspace for running the RAW DATA → VALIDATION → CLEANING → PROCESSED DATA flow on a dataset, and for seeing the resulting cleaned dataset and the record of what was done.
  • Primary actions: run the cleaning pipeline; inspect the cleaned dataset; open the transformation record.
  • Supporting actions: open the validation report; return to the dataset listing.
  • Domain entities: raw dataset, cleaned dataset, validation report, transformation record.
  • Component responsibilities: four-stage pipeline strip with PEMBERSIHAN or PROCESSED as the current stage; determinate tangerine progress bar while the pipeline runs; summary of transformations applied; link to the cleaned dataset and to the data issues list.
  • States: empty — no dataset selected or no dataset exists; the page states that no dataset is available. loading — the pipeline is running, shown as a determinate progress bar. success — cleaned dataset produced and the transformation record is available. error — the pipeline failed; the failure is stated honestly and the raw dataset is left intact. recovery — the user can re-run cleaning after addressing the reported issues.

Data Issues

  • Purpose and information/state: the report of data problems found in the dataset and the record of important transformations that were performed, so that nothing is deleted silently.
  • Primary actions: read each issue and each transformation note; follow the action link on an issue.
  • Supporting actions: open the validation report; return to the dataset listing.
  • Domain entities: data issue (category, severity, affected-row count, human explanation), transformation record (what was done, to which column, affecting how many values).
  • Component responsibilities: ruled ledger of issues and transformations — severity chip, monospace affected-row count, human sentence such as "Kolom tanggal memiliki 12 nilai yang tidak dapat dibaca.", right-side action link; no coloured alert cards.
  • States: empty — no issues have been recorded because no validation or cleaning has run, or because the dataset is clean; the page states which of the two applies. loading — issues are being read. success — issues and transformation notes listed with real counts. error — issues could not be read. recovery — re-run validation or cleaning.
Page 8 of 28

3. Functional Requirements

Page 9 of 28

Phase 1 — Project Foundation

FR-1. Project scaffold structure — explicit As a developer, I should have a project rooted at jayapura-retail-intelligence/ containing frontend/, backend/, analytics/, ai/, data/raw/, data/processed/, data/external/, reports/, tests/, docs/, README.md, .gitignore, and .env.example, so that the project has a clean and professional foundation.

  • Observable acceptance: every listed directory and file exists at the stated path.
  • Continuation: the remaining Phase 1 requirements build on this structure.

FR-2. Frontend stack — explicit As a developer, I should have the frontend built with React, TypeScript, and Vite, so that the interface is developed on the specified stack.

  • Observable acceptance: the frontend builds successfully with the specified stack.

FR-3. Backend stack — explicit As a developer, I should have the backend built with Python and FastAPI, so that the API is developed on the specified stack.

  • Observable acceptance: the Python backend runs.

FR-4. Data stack — explicit As a developer, I should have the data layer built with Pandas, NumPy, and openpyxl, so that dataset reading, validation, and cleaning use the specified libraries.

  • Observable acceptance: Pandas, NumPy, and openpyxl are the libraries used for dataset handling.

FR-5. AI not activated — explicit As a developer, I should have no AI activated in this master prompt, so that the delivery stays within the DATA FIRST boundary.

  • Observable acceptance: no AI agent, chatbot, or AI interpretation is present or invoked.

FR-6. Database deferred — explicit As a developer, I should not introduce a database at this stage unless the architecture genuinely requires one, so that the delivery stays minimal.

  • Observable acceptance: no business database is present unless a genuine architectural need is documented.

FR-7. Extensible code structure — explicit As a developer, I should have a code structure that is easy to extend, so that later master prompts can build on it.

  • Observable acceptance: modules are separated by responsibility across frontend, backend, analytics, and ai.

FR-8. Phase 1 exclusions — explicit As a developer, I should not build login, authentication, AI agent, chatbot, web scraping, real market data, dummy sales data, dummy revenue, dummy customer, dummy inventory, business database, API key, or fake business intelligence, so that the foundation contains no prohibited capability.

  • Observable acceptance: none of the listed items exists in the delivered project.
Page 10 of 28

Phase 2 — Data Workspace

FR-9. CSV dataset upload — explicit As an Analis Data Retail, I should be able to upload a business dataset in CSV format, so that my own sales data can enter the workspace.

  • Trigger/input: a CSV file selected or dropped on the Upload CSV page.
  • Observable result: a dataset record is created and appears in the dataset listing with its real measured row and column counts.
  • Access state: no login required.
  • Material failure/recovery: if the file is not a CSV or cannot be parsed as CSV, the upload is rejected with a message stating the reason, and the user can select a different file.
  • Continuation: the user proceeds to the dataset's status, metadata, preview, and validation status.

FR-10. XLSX dataset upload — explicit As an Analis Data Retail, I should be able to upload a business dataset in XLSX format, so that my own sales data can enter the workspace.

  • Trigger/input: an XLSX file selected or dropped on the Upload XLSX page.
  • Observable result: a dataset record is created and appears in the dataset listing with its real measured row and column counts.
  • Access state: no login required.
  • Material failure/recovery: if the file is not an XLSX or cannot be parsed as XLSX, the upload is rejected with a message stating the reason, and the user can select a different file.
  • Continuation: the user proceeds to the dataset's status, metadata, preview, and validation status.

FR-11. File type validation on upload — explicit As an Analis Data Retail, I should have the system validate the file type before accepting a file as a business dataset, so that arbitrary files are not treated as business data.

  • Trigger/input: any file submitted to an upload page.
  • Observable result: only CSV and XLSX files are accepted as business datasets; anything else is rejected with a stated reason.
  • Material failure/recovery: rejection is explicit and the user can retry with a valid file.
  • Continuation: accepted files proceed to dataset creation.

FR-12. Clean and safe upload endpoints — explicit As a developer, I should provide clean and safe backend endpoints for the upload feature where the upload requires a backend API, so that ingestion is handled properly.

  • Observable result: upload endpoints exist and enforce file type validation.
  • Continuation: the endpoints support the upload pages.

FR-13. Dataset listing — explicit As an Analis Data Retail, I should see a listing of the datasets available in the Data Workspace, so that I know what data I have brought in.

  • Observable result: each dataset appears with its identity and current state, using real measured values.
  • Access state: no login required.
  • Material failure/recovery: if the listing cannot be read, the failure is stated and the read can be retried.
  • Continuation: the user opens a dataset's status, metadata, preview, or validation status.

FR-14. Dataset status — explicit As an Analis Data Retail, I should see the processing and validation status of a dataset, so that I know where it stands in the pipeline.

  • Observable result: the dataset's current stage in RAW → VALIDASI → PEMBERSIHAN → PROCESSED is shown truthfully.
  • Material failure/recovery: if status cannot be read, the failure is stated and the read can be retried.
  • Continuation: the user runs validation or cleaning, or opens the validation report.

FR-15. Dataset metadata — explicit As an Analis Data Retail, I should see metadata for a dataset, including its dataset type and data structure, so that I understand what I uploaded.

  • Observable result: metadata and the column inventory are shown, with each column's role marked as REQUIRED, OPTIONAL, or DERIVED.
  • Material failure/recovery: if metadata cannot be read, the failure is stated and the read can be retried.
  • Continuation: the user opens the validation report or the preview.

FR-16. Dataset preview — explicit As an Analis Data Retail, I should see a preview of the actual rows of a dataset, so that I can confirm the file contains what I expect.

  • Observable result: real rows from the user's file are rendered; no sample rows are ever shown.
  • Material failure/recovery: if the preview cannot be read, the failure is stated and the read can be retried.
  • Continuation: the user proceeds to validation status or cleaning.

FR-17. Data validation status — explicit As an Analis Data Retail, I should see the validation status of a dataset against the data contract, so that I know whether it is usable.

  • Observable result: a validation summary with real affected-row counts and severity markers.
  • Material failure/recovery: if validation cannot be run or results cannot be read, the failure is stated and validation can be re-run.
  • Continuation: the user opens the validation report or the data issues list.

FR-18. Honest empty state when no dataset exists — explicit As a Pemilik/Pengelola Bisnis Retail, I should see an honest empty state stating that no dataset is available when no dataset has been uploaded, so that I am never shown fabricated business information.

  • Observable result: the workspace states that the dataset is not yet available. It does not show Rp0 as business data, "100 transaksi", a fake chart, a sample product, or a fake KPI.
  • Continuation: the user is offered the upload entry.

FR-19. No sample data — explicit As a Pemilik/Pengelola Bisnis Retail, I should never see example data in the workspace, so that everything I see is my own.

  • Observable result: no sample dataset, sample product, or sample transaction is present anywhere.
Page 11 of 28

Phase 3 — Data Cleaning & Validation

FR-20. Cleaning pipeline flow — explicit As an Analis Data Retail, I should have the data flow RAW DATA → VALIDATION → CLEANING → PROCESSED DATA implemented with Python/Pandas, so that my raw file becomes a processed dataset through a defined sequence.

  • Observable result: the four stages are visible and the dataset moves through them.
  • Material failure/recovery: if a stage fails, the failure is stated and the raw dataset is left intact.
  • Continuation: the user obtains a cleaned dataset and a validation report.

FR-21. No business analysis in Phase 3 — explicit As a developer, I should not perform business analysis in Phase 3, so that cleaning stays separate from analysis.

  • Observable result: no sales, market, customer, technology, inventory, accounting, or BI analysis is produced.

FR-22. Detection of missing required columns — explicit As an Analis Data Retail, I should have the system detect that a required column is missing, so that I know the dataset cannot satisfy the contract.

  • Observable result: the missing required column is reported as a finding with a plain sentence.

FR-23. Detection of mismatched column names — explicit As an Analis Data Retail, I should have the system detect column names that do not match the contract, so that I can correct my headers.

  • Observable result: the mismatched column name is reported as a finding.

FR-24. Detection of invalid dates — explicit As an Analis Data Retail, I should have the system detect invalid date values, so that I know which rows cannot be read as dates.

  • Observable result: a finding states the column and the number of unreadable values, for example "Kolom tanggal memiliki 12 nilai yang tidak dapat dibaca."

FR-25. Detection of numbers read as text — explicit As an Analis Data Retail, I should have the system detect numeric values that were read as text, so that I can fix the source formatting.

  • Observable result: a finding states the column and the number of affected values.

FR-26. Detection of empty unit — explicit As an Analis Data Retail, I should have the system detect empty unit values, so that I know where quantity is missing.

  • Observable result: a finding states the column and the number of empty values.

FR-27. Detection of empty harga — explicit As an Analis Data Retail, I should have the system detect empty harga values, so that I know where price is missing.

  • Observable result: a finding states the column and the number of empty values.

FR-28. Detection of empty revenue — explicit As an Analis Data Retail, I should have the system detect empty revenue values, so that I know where revenue is missing.

  • Observable result: a finding states the column and the number of empty values.

FR-29. Detection of duplicate transactions — explicit As an Analis Data Retail, I should have the system detect duplicate transactions, so that I know where the same transaction appears more than once.

  • Observable result: a finding states the number of duplicated transactions.

FR-30. Detection of suspicious negative values — explicit As an Analis Data Retail, I should have the system detect suspicious negative values, so that I can review them before they affect anything downstream.

  • Observable result: a finding states the column and the number of suspicious negative values.

FR-31. Detection of differing date formats — explicit As an Analis Data Retail, I should have the system detect that date formats differ within a column, so that I know the column is not uniform.

  • Observable result: a finding states the column and the number of values in the differing format.

FR-32. Detection of whitespace — explicit As an Analis Data Retail, I should have the system detect whitespace problems in values, so that inconsistent text does not silently split categories.

  • Observable result: a finding states the column and the number of affected values.

FR-33. Detection of inconsistent category names — explicit As an Analis Data Retail, I should have the system detect inconsistent kategori names, so that the same category is not counted as several.

  • Observable result: a finding states the affected category values and their count.

FR-34. Detection of inconsistent product names — explicit As an Analis Data Retail, I should have the system detect inconsistent produk names, so that the same product is not counted as several.

  • Observable result: a finding states the affected product values and their count.

FR-35. Raw dataset, cleaned dataset, and validation report concepts — explicit As an Analis Data Retail, I should have a raw dataset, a cleaned dataset, and a validation report as distinct concepts, so that I can always compare what I uploaded with what was produced.

  • Observable result: all three exist and are separately inspectable.

FR-36. Explain data problems to the user — explicit As an Analis Data Retail, I should have the system explain data problems to me in plain language, so that I understand what is wrong with my file.

  • Observable result: each problem is stated as a human sentence naming the column and the affected count.

FR-37. No silent deletion — explicit As an Analis Data Retail, I should never have data deleted silently, so that I retain control over my own file.

  • Observable result: no row or value is removed without a corresponding recorded explanation.

FR-38. Record important transformations — explicit As an Analis Data Retail, I should have the system record what it did for every important transformation, so that the cleaned dataset is auditable.

  • Observable result: each important transformation is stored with what was done, to which column, and affecting how many values.
Page 12 of 28

Data Contract

FR-39. Data contract definition — explicit As a developer, I should have a clear data contract defining dataset type, required columns, optional columns, data types, and validation rules, so that validation and cleaning have a single authoritative definition.

  • Observable result: the contract is defined and is the source used by validation.

FR-40. Sales dataset canonical fields — explicit As a developer, I should define canonical fields for the sales dataset as tanggal, transaksi_id, pos, produk, kategori, brand, unit, harga, diskon, revenue, and customer_id, so that sales data has a stable vocabulary.

  • Observable result: these canonical field names are the ones the contract uses.

FR-41. Required, optional, and derived distinction — explicit As a developer, I should distinguish REQUIRED, OPTIONAL, and DERIVED fields, and not force every field to be present when it is not needed, so that the contract fits real files.

  • Observable result: each canonical field carries one of the three roles, and a dataset missing an optional field is not rejected for that reason.

FR-42. Revenue as a derived field — explicit As an Analis Data Retail, I should have revenue treated as DERIVED when unit and harga are available, so that revenue can be computed rather than required in my file.

  • Observable result: when revenue is absent but unit and harga are present, revenue is computed and the column is tagged DERIVED.
  • Continuation: the derived column is visible as derived in metadata and preview.
Page 13 of 28

Aturan Angka

FR-43. All numbers computed from the dataset — explicit As a Pemilik/Pengelola Bisnis Retail, I should have every number that comes from a dataset genuinely computed from that dataset, so that I can trust it.

  • Observable result: every displayed count, size, and percentage traces to a real measurement of a real file.

FR-44. No demo numbers, random data, or hardcoded business numbers — explicit As a Pemilik/Pengelola Bisnis Retail, I should never see a number created for a demo, generated randomly, or hardcoded as business data, so that nothing in the workspace is fabricated.

  • Observable result: no random generation and no hardcoded business figure exists in the delivered system.
Page 14 of 28

UI

FR-45. Indonesian interface language — explicit As an Analis Data Retail, I should have the main interface in Indonesian, so that I can work in the language of my business.

  • Observable acceptance: primary interface copy is in Indonesian.

FR-46. Indonesian number and date formats and WIB timezone — explicit As an Analis Data Retail, I should see Rupiah formatted with Rp, dates formatted as DD/MM/YYYY, and times in WIB, so that values match local convention.

  • Observable acceptance: currency uses Rp, dates use DD/MM/YYYY, and timezone is WIB.

FR-47. No business numbers before a dataset exists — explicit As a Pemilik/Pengelola Bisnis Retail, I should see no business numbers at all while no dataset is available, so that the absence is stated rather than filled with a placeholder.

  • Observable acceptance: no Rp0, no transaction count, no chart, no sample product, and no KPI is rendered before a real dataset exists.

FR-48. Simple, professional, extensible UI — explicit As a developer, I should have a UI that is simple, professional, and easy to extend, so that later master prompts can build on it.

  • Observable acceptance: the interface follows the stated visual system and its components are reusable.
Page 15 of 28

Testing

FR-49. File validation tests — explicit As a developer, I should have basic tests for file validation, so that rejected file types are proven to be rejected.

  • Observable acceptance: tests exist and pass or fail honestly.

FR-50. Dataset validation tests — explicit As a developer, I should have basic tests for dataset validation, so that contract checks are proven.

  • Observable acceptance: tests exist and report honestly.

FR-51. Column validation tests — explicit As a developer, I should have basic tests for column validation, so that missing and mismatched columns are proven to be detected.

  • Observable acceptance: tests exist and report honestly.

FR-52. Data type validation tests — explicit As a developer, I should have basic tests for data type validation, so that type problems are proven to be detected.

  • Observable acceptance: tests exist and report honestly.

FR-53. Cleaning pipeline tests — explicit As a developer, I should have basic tests for the cleaning pipeline, so that the RAW → VALIDATION → CLEANING → PROCESSED flow is proven.

  • Observable acceptance: tests exist and report honestly.

FR-54. Empty dataset state tests — explicit As a developer, I should have basic tests for the empty dataset state, so that the honest empty state is proven.

  • Observable acceptance: tests exist and report honestly.

FR-55. Backend runs, frontend builds, no serious lint errors — explicit As a developer, I should have the Python backend runnable, the frontend buildable, and no serious lint errors, so that the project is in a working state.

  • Observable acceptance: backend starts, frontend builds, and lint reports no serious errors.

FR-56. Honest test reporting — explicit As a developer, I should report test results honestly without inventing them, and fix failures where possible, so that the reported state matches reality.

  • Observable acceptance: reported results match actual runs; failures are disclosed.
Page 16 of 28

README

FR-57. README content — explicit As a developer, I should have the README updated with the project purpose, architecture, folder structure, data flow, data contract, roadmap, the DATA FIRST principle, the deterministic analytics principle, and the status of completed phases, so that the project is documented.

  • Observable acceptance: every listed item is present in the README.

Final Deliverable

FR-58. Master Prompt 1 final report — explicit As a developer, I should deliver the final structure, the list of important files created, the architecture explanation, the Data Workspace explanation, the cleaning pipeline explanation, the data contract explanation, the list of endpoints created, the list of tests, the lint/build/test results, and the list of things NOT yet built, so that the work is fully accounted for.

  • Observable acceptance: all ten items are provided.

FR-59. Stop after Master Prompt 1 — explicit As a developer, I should stop after Master Prompt 1 and not begin Master Prompt 2 before the next instruction, so that the delivery boundary is respected.

  • Observable acceptance: no Master Prompt 2 module is built.

4. User Personas

Page 17 of 28

Analis Data Retail

Product context. This persona prepares the shop's own data. They export sales records from a POS or spreadsheet, and the file they hold is rarely clean: headers drift between exports, dates arrive in more than one format, quantities are blank, and the same product is spelled three ways. They are the person who has to make that file trustworthy before anyone else looks at it.

Primary goal. Bring a CSV or XLSX business dataset into the workspace, understand exactly what is wrong with it, and produce a cleaned dataset with a validation report and a record of what was changed — without any business analysis being performed yet.

Distinct accepted responsibilities.

  • Upload a CSV dataset and upload an XLSX dataset, each through its own workspace.
  • Read the dataset listing to see what has been brought in.
  • Read a dataset's status to know where it sits in RAW → VALIDASI → PEMBERSIHAN → PROCESSED.
  • Read a dataset's metadata, including its dataset type and its column inventory with REQUIRED / OPTIONAL / DERIVED roles.
  • Read a preview of the actual rows of the dataset.
  • Read the validation status and the full validation report against the data contract.
  • Run the cleaning pipeline and obtain a cleaned dataset.
  • Read the data issues list and the transformation record.

Relevant inputs and decisions. The input is the user's own file. The decisions are: whether the file is the right one (confirmed by preview), whether the detected problems are acceptable or must be fixed at source, and whether to proceed to cleaning. The persona decides when a dataset is good enough to move forward.

Interactions with other accepted participants. This persona produces the cleaned dataset and the validation report that the Pemilik/Pengelola Bisnis Retail reads to judge whether the data is fit to use. The analyst's work is the precondition for the owner's confidence.

Observable success. A cleaned dataset exists, a validation report explains the problems found, and a transformation record shows what was done — with every count traceable to the uploaded file.

What makes this role different. This persona operates on the data itself: they upload, inspect, validate, and clean. Their work is the pipeline.

Page 18 of 28

Pemilik/Pengelola Bisnis Retail

Product context. This persona owns or manages an electronics/gadget retail business in Jayapura or the surrounding area. They are not going to clean files themselves, but they are the person who will eventually act on numbers, and they have been burned before by dashboards that showed impressive figures that turned out to be examples.

Primary goal. Be certain that any number the system will eventually show comes from a real dataset rather than a demo, and that the system says plainly when there is no data.

Distinct accepted responsibilities.

  • Read the dataset listing and each dataset's status to see what data actually exists.
  • Read the validation report and the data issues list to judge whether the data is fit to use.
  • Receive an honest empty state when no dataset is available, rather than a placeholder figure.

Relevant inputs and decisions. The input is the analyst's uploaded dataset and its validation outcome. The decision is whether the data is trustworthy enough to be used as the basis for later analysis.

Interactions with other accepted participants. This persona depends on the Analis Data Retail's upload and cleaning work. They consume the validation report and the data issues list; they do not produce them.

Observable success. The workspace shows either a real dataset with real measured counts, or a plain statement that no dataset is available. At no point does it show Rp0 as business data, a transaction count, a chart, a sample product, or a KPI that was not computed from a real file.

What makes this role different. This persona does not operate on the data; they audit the system's honesty about the data. Their success condition is the absence of fabrication, which is a different kind of requirement from the analyst's success condition of a clean dataset.

Page 19 of 28

5. Core User Flows

Flow 1 — Analis Data Retail uploads a CSV dataset and inspects it

  1. The analyst opens the workspace. No login is required. The Landing page is shown.
  2. Because no dataset exists, Landing displays the honest empty state: a flush-left statement that no dataset is available, a ruled horizontal rule, and one sentence explaining that a CSV or XLSX file can be uploaded and that all numbers will be computed from the user's data rather than from examples. No Rp0, no transaction count, no chart, no sample product, and no KPI is rendered.
  3. The analyst chooses the primary upload entry and arrives at Upload CSV.
  4. The analyst selects or drops a CSV file onto the dashed dropzone. While the file is dragged over it, the dropzone border becomes 2px dashed tangerine.
  5. The system validates the file type. If the file is not a CSV, the upload is rejected and the page states the reason; the analyst selects a different file and resubmits. If the file is a CSV, a determinate tangerine progress bar shows the upload.
  6. On success, a dataset record is created. The analyst is offered the dataset's status, metadata, and preview.
  7. The analyst opens Datasets and sees the new dataset listed in a ruled ledger with its real measured row and column counts, right-aligned in tabular numerals.
  8. The analyst opens Dataset Metadata and reads the dataset type and the column inventory, with each column marked REQUIRED, OPTIONAL, or DERIVED. If revenue was absent but unit and harga were present, the revenue column carries the acid-lime DERIVED chip.
  9. The analyst opens Dataset Preview and reads the actual rows of the file in a horizontally scrollable ruled table with a frozen first column. No sample rows appear.
  10. Failure/recovery: if the listing, metadata, or preview cannot be read, the page states the failure and the analyst retries the read.
  11. Continuation: the analyst proceeds to validation.

Flow 2 — Analis Data Retail uploads an XLSX dataset

  1. The analyst opens the workspace. No login is required.
  2. The analyst navigates to Upload XLSX.
  3. The analyst selects or drops an XLSX file onto the dashed dropzone.
  4. The system validates the file type. A non-XLSX file is rejected with a stated reason and the analyst retries. A valid XLSX file uploads behind a determinate tangerine progress bar.
  5. On success, a dataset record is created and appears in Datasets with real measured counts.
  6. Failure/recovery: if the file cannot be parsed as XLSX, the failure is stated and the analyst selects a different file.
  7. Continuation: the analyst proceeds to validation, exactly as in Flow 1.
Page 20 of 28

Flow 3 — Analis Data Retail validates a dataset against the data contract

  1. The analyst opens Dataset Status for the dataset. The four-stage pipeline strip shows RAW as the current stage, filled tangerine, with VALIDASI, PEMBERSIHAN, and PROCESSED hairline-outlined.
  2. The analyst runs validation.
  3. The system applies the data contract: dataset type, required columns, optional columns, data types, and validation rules.
  4. The analyst opens Validation Status and sees the summary, with VALIDASI now the current stage in the pipeline strip. Severity chips read tangerine Perlu tindakan, lime Bersih, or muted Info, and affected-row counts are set in tabular numerals.
  5. The analyst opens Validation Report and reads the ruled ledger of findings. Each finding is one row: a severity chip, a monospace affected-row count, a human sentence, and a right-side action link. A date problem reads, for example, "Kolom tanggal memiliki 12 nilai yang tidak dapat dibaca." The report also shows which required columns, optional columns, data types, and validation rules were applied.
  6. The analyst opens Data Issues to read the full list of problems, including missing required columns, mismatched column names, invalid dates, numbers read as text, empty unit, empty harga, empty revenue, duplicate transactions, suspicious negative values, differing date formats, whitespace, inconsistent kategori names, and inconsistent produk names.
  7. Failure/recovery: if validation cannot be run or the report cannot be produced, the failure is stated and the analyst re-runs validation.
  8. Continuation: the analyst decides whether to fix the source file or proceed to cleaning.

Flow 4 — Analis Data Retail runs cleaning and obtains a processed dataset

  1. The analyst opens Data Cleaning for a dataset that has been validated.
  2. The analyst runs the cleaning pipeline. The pipeline strip shows PEMBERSIHAN as the current stage, and a determinate tangerine progress bar shows the run.
  3. The system executes RAW DATA → VALIDATION → CLEANING → PROCESSED DATA using Python/Pandas. No business analysis is performed.
  4. No data is deleted silently. Every important transformation is recorded with what was done, to which column, and affecting how many values.
  5. On success, the pipeline strip shows PROCESSED as the current stage. A cleaned dataset exists alongside the raw dataset, and the transformation record is available.
  6. The analyst opens Data Issues and reads both the problem list and the transformation notes, each as a ruled row with a severity chip, a monospace affected-row count, and a human sentence.
  7. The analyst opens Dataset Preview and switches to the cleaned preview to compare it with the raw preview.
  8. Failure/recovery: if the pipeline fails, the failure is stated honestly and the raw dataset is left intact. The analyst addresses the reported issues and re-runs cleaning.
  9. Continuation: the cleaned dataset and validation report are now available for the owner to review. Business analysis is not performed in this delivery.

Flow 5 — Pemilik/Pengelola Bisnis Retail confirms that no fabricated data is shown

  1. The owner opens the workspace. No login is required. The Landing page is shown.
  2. Because no dataset exists, the owner sees the honest empty state: a plain statement that no dataset is available. There is no Rp0 presented as business data, no "100 transaksi", no fake chart, no sample product, and no fake KPI.
  3. The owner opens Datasets and sees the same honest empty state, with the upload entry offered.
  4. Continuation: the owner waits for the analyst to upload a dataset, or asks the analyst to do so.
Page 21 of 28

Flow 6 — Pemilik/Pengelola Bisnis Retail reviews whether the data is fit to use

  1. The analyst has uploaded and validated a dataset. The owner opens the workspace. No login is required.
  2. The owner opens Datasets and sees the dataset listed with its real measured row and column counts.
  3. The owner opens Dataset Status and reads the current stage in the RAW → VALIDASI → PEMBERSIHAN → PROCESSED strip.
  4. The owner opens Dataset Metadata and reads the dataset type and the column inventory, noting which fields are REQUIRED, OPTIONAL, or DERIVED, and which columns carry the DERIVED chip.
  5. The owner opens Validation Status and reads the summary with its severity chips and real affected-row counts.
  6. The owner opens Validation Report and reads each finding as a plain sentence naming the column and the affected count.
  7. The owner opens Data Issues and reads the problem list and the transformation record, confirming that nothing was deleted silently.
  8. Failure/recovery: if a report cannot be read, the page states the failure and the owner retries.
  9. Continuation: the owner judges whether the data is trustworthy enough to be used as the basis for later analysis. No business analysis is performed in this delivery.
Page 22 of 28

6. Visuals Colors and Theme

Muse and headline. Rasmus Andersson — systematic product craft with an opinion. The register is custody, not excitement: this tool must read as a precision instrument that will not lie about the user's data.

Mode. Dark mode only.

Colour tokens by role.

RoleTokenValue
Page ground--bg#101113
Panel / table surface--surface#16181B
Ink (primary text)--text#EDEAE4
Signal (primary)--primary#FF6B2C
Valid / derived status--accent#C6F24E
Muted (labels, metadata keys, timestamps, empty-state copy)--muted#8A8A85
Hairline border--hairlinergba(237,234,228,0.10)

#101113 is the page ground — warm graphite, never pure black. #16181B is the panel and table surface, separated from the ground by 1px hairlines. #EDEAE4 is warm off-white ink at 14.5:1 on the ground. #FF6B2C is THE signal: primary buttons, focus rings, the active upload dropzone border, the current nav underline, and the single Perlu tindakan severity marker — used on roughly 3% of pixels. #C6F24E is a rarer second status colour reserved exclusively for valid / bersih state chips and the DERIVED field badge, never decorative. #8A8A85 carries labels, metadata keys, timestamps, and honest empty-state copy.

Ratio. Approximately 78% graphite ground, 15% panel surface, 4% ink, 3% tangerine and lime combined.

No blue. There is no blue, indigo, or violet anywhere in the system, including links. Links are ink with a tangerine underline.

Typography.

  • Headings: Space Grotesk, weights 500–700, tight tracking (−0.02em to −0.035em), sentence case, with occasional all-caps micro-labels at 11px / 0.14em letterspacing. Headlines are set large and flush-left, never centred.
  • Body: Inter Tight.
  • Data numerals: JetBrains Mono at 500 weight with tabular figures for every count, row index, byte size, and percentage, so columns align optically across the workspace.
  • Type scale: 1.25 modular on a 4/8-pt grid — display 40px mobile → 72px desktop (clamp), section head 24/28, panel title 18, body 15, table cell 13.5, micro-label 11 uppercase.
  • Line height: 1.05 on display, 1.55 on body.
  • Numerals: always tabular-nums.

Shape language. Rectilinear and honest. 6px radius on controls, 8px on panels, 3px on chips. No pill buttons, no blobs, no soft offset shadows. Separation comes from 1px hairlines and one-step surface lifts, never from elevation blur. The upload dropzone is a dashed 1px rectangle, becoming 2px dashed tangerine only while dragging over it. Table rows are ruled, not carded; row hover is a flat surface fill, not a lift.

Spacing rhythm. 4/8-pt grid throughout. Main column is a strict 12-column grid, max-width 1440px, 24px gutters mobile / 32px desktop.

Layout. A fixed left rail — 72px icon rail at 375–767px, 232px labelled rail at ≥1024px — carrying the workspace sections: Ringkasan Data, Unggah Dataset, Daftar Dataset, Validasi & Pembersihan, Kontrak Data. The dataset detail view is a two-pane split: metadata and validation report on the left (5 columns, sticky), raw/cleaned preview table on the right (7 columns, horizontally scrollable with a frozen first column). Every panel header is a ruled bar with the title left and a status chip right. Nothing is centred; everything is flush-left on the grid.

Imagery style. The interface is the imagery. No photography, no illustration, no 3D. Visual interest is carried by tabular numeral blocks, a monospace schema listing of the data contract, hairline-grid diagrams showing the four-stage pipeline as ruled boxes with arrows, and a sparkline-free distribution strip of column fill rates rendered as a row of 1px vertical bars. The empty state is a single centred 1px-ruled rectangle with muted copy — an absent shelf, not a cartoon.

Accessibility. Text and controls stay whole at 375px, 768px, and 1280px: headlines, labels, numbers, and controls remain entirely inside the viewport and their container, wrapping or scaling to fit, and no other element covers any part of them. Decoration may be cropped or bled off an edge; readable text and controls may not.

Page 23 of 28

7. Signature Design Concept

The honest absence. The first screen is the workspace itself, opened on Ringkasan Data with zero datasets loaded. There is no marketing hero.

The composition: a full-height graphite field (#101113), a 72px icon rail hard against the left edge, and in the main column a flush-left display line in Space Grotesk at 40px mobile → 72px desktop reading "Belum ada dataset." set on two lines, the second line in muted #8A8A85. Directly beneath it, a single 1px-ruled horizontal rule spanning the full content width. Beneath that, one sentence in body copy: "Unggah file CSV atau XLSX untuk memulai. Semua angka akan dihitung dari data Anda, bukan dari contoh." Beneath that, one solid tangerine rectangle button (Unggah Dataset) 44px tall, flush-left, and to its right a dashed 1px outlined secondary target.

The right two-thirds of the viewport stays empty graphite. That negative space is deliberate: it reads as an instrument waiting for a specimen. No KPI tiles, no Rp0, no charts, no illustration, no gradient.

This concept recomposes only accepted content and controls — the honest empty state, the upload entry, and the workspace shell. It introduces no new behavior, page, or destination.

Page 24 of 28

8. Interaction Model & Motion Direction

Interaction Model: Static Motion Tempo: restrained Hero Dimensionality: flat

Landing Hero Motion Brief

  • Focal subject: the oversized flush-left Space Grotesk statement "Belum ada dataset." over an empty ruled graphite field, with the right two-thirds of the viewport deliberately blank.
  • Input → transformation → outcome thesis: the user's only input on this screen is choosing the upload entry. The transformation is not visual — it is the arrival of a real file, which replaces the empty state with real measured counts. The outcome is that the emptiness is never filled with a fabricated figure; it is filled only by the user's own data.
  • Motion vocabulary: functional only, 120–180ms, cubic-bezier(0.2, 0, 0, 1). Upload progress is a determinate tangerine bar, not a spinner. Validation report rows stagger in at 40ms intervals on first paint and then never animate again. Table row hover is an instant surface fill (90ms). Focus rings appear in 0ms. No parallax, no scroll reveals, no bounce, no counters ticking up.
  • Composed first frame: graphite ground, 72px icon rail at the left edge, the two-line display statement flush-left, the ruled horizontal rule, the single body sentence, the solid tangerine button, the dashed secondary target, and empty graphite to the right.
  • Reduced-motion state: with prefers-reduced-motion, every transition collapses to 0ms and the validation-report stagger becomes a single paint. The first frame is unchanged.

No 3D or WebGL scene is required or requested for this delivery.

Page 25 of 28

9. Non-Functional Requirements

NFR-1. Deterministic computation before interpretation — explicit Python/Pandas must compute numbers and statistics first; AI interpretation comes second and must never invent numbers. In this delivery AI is not activated at all. Rationale: this is the governing product principle.

NFR-2. No fabricated business content — explicit No demo numbers, no random data, and no hardcoded business figures may exist anywhere in the system. Rationale: the product's entire value is that the user can trust what it says about their own data.

NFR-3. Honest empty state — explicit When no dataset exists, the system must state that the dataset is not yet available. It must not render Rp0 as business data, "100 transaksi", a fake chart, a sample product, or a fake KPI. Rationale: an empty state that lies destroys the custody promise.

NFR-4. File type validation — explicit Uploads must validate file type; arbitrary files must not be accepted as business datasets. Rationale: the workspace must not ingest non-business files.

NFR-5. No silent data deletion — explicit Data must never be removed silently; problems must be explained to the user. Rationale: the user retains custody of their file.

NFR-6. Transformation traceability — explicit Every important transformation must be recorded with what was done. Rationale: the cleaned dataset must be auditable against the raw dataset.

NFR-7. No forced fields — explicit The data contract must not force every field to be present when it is not needed. Rationale: real exports vary, and a contract that rejects valid files is useless.

NFR-8. Honest test reporting — explicit Test results must be reported honestly and never invented; failures must be disclosed and fixed where possible. Rationale: the project's credibility depends on truthful reporting.

NFR-9. Indonesian locale conventions — explicit The main interface is in Indonesian; currency uses Rp; dates use DD/MM/YYYY; timezone is WIB. Rationale: the audience operates in Indonesian retail convention.

NFR-10. Runnable backend, buildable frontend, clean lint — explicit The Python backend must run, the frontend must build, and there must be no serious lint errors. Rationale: the project must be in a working state.

NFR-11. No login or authentication — explicit No login and no authentication may be built in this delivery. Rationale: explicitly excluded by the user for this master prompt.

NFR-12. No AI, no database, no API key — explicit AI is not activated; a database is not required unless the architecture genuinely requires one; no API key is created. Rationale: explicitly excluded or deferred by the user for this master prompt.

NFR-13. Readable text and controls at every viewport — required_inference Headlines, labels, numbers, and controls must remain entirely inside the viewport and their container at 375px, 768px, and 1280px, wrapping or scaling to fit, with no other element covering them. Rationale: required to make the accepted interface usable at the stated breakpoints.

Page 26 of 28

10. Tech Stack

All technology choices below are explicit user requirements.

Frontend

  • React
  • TypeScript
  • Vite

Backend

  • Python
  • FastAPI

Data

  • Pandas
  • NumPy
  • openpyxl

AI

  • Not activated in this master prompt.

Database

  • Not required at this stage unless the architecture genuinely requires one.

Project structure

jayapura-retail-intelligence/
├── frontend/
├── backend/
├── analytics/
├── ai/
├── data/
│   ├── raw/
│   ├── processed/
│  \x20\xe2\x94\x94── external/
├── reports/
├── tests/
├── docs/
├── README.md
├── .gitignore
\xe2\x94\x94── .env.example

Containerization and orchestration

  • [Default — not specified by user] Docker and docker-compose for local development and running the frontend, backend, and analytics together.
  • [Default — not specified by user] Kubernetes is not required for this delivery.
Page 27 of 28

11. Assumptions and Constraints

Assumptions

  • A-1 [Assumption] The user's CSV and XLSX files are exports from their own POS or spreadsheet tooling and are not encrypted or password-protected.
  • A-2 [Assumption] A single uploaded file corresponds to a single dataset.
  • A-3 [Assumption] The sales dataset is the only dataset type defined in this delivery; the data contract mechanism is general enough to add further dataset types later.
  • A-4 [Assumption] The workspace is used by a small number of people in the same business, which is why no login and no differentiated permissions are needed in this delivery.
  • A-5 [Assumption] The ai/ directory exists as structure only and contains no active AI behavior.

Constraints

  • C-1 Do not build login, authentication, AI agent, chatbot, web scraping, real market data, dummy sales data, dummy revenue, dummy customer, dummy inventory, business database, API key, or fake business intelligence.
  • C-2 Do not create example data or fake data; do not create numbers for a demo; do not use random data; do not use hardcoded numbers as business data.
  • C-3 Do not display Rp0 as business data, "100 transaksi", a fake chart, a sample product, or a fake KPI when no dataset is available.
  • C-4 Do not accept arbitrary files as business datasets; use file type validation.
  • C-5 Do not perform business analysis first in Phase 3.
  • C-6 Do not delete data silently; explain data problems to the user.
  • C-7 Do not force every field to be present when it is not needed.
  • C-8 Do not invent test results; report failures honestly.
  • C-9 Do not build a complete Sales Analytics Engine, Market Intelligence, Customer Intelligence, Technology Intelligence, Inventory Analytics, Accounting, AI Business Analyst, Report Generator, or Final Dashboard — all of these will be done in the following master prompts.
  • C-10 Do not continue into modules that are not part of Master Prompt 1.
  • C-11 After Master Prompt 1 is complete, stop and do not work on Master Prompt 2 before the next instruction.
  • C-12 AI is not activated in this master prompt.
  • C-13 A database is not required at this stage unless the architecture genuinely requires one.

Future requirements (not current, not in scope for this delivery)

  • Sales Analytics Engine (complete)
  • Market Intelligence
  • Customer Intelligence
  • Technology Intelligence
  • Inventory Analytics
  • Accounting
  • AI Business Analyst
  • Report Generator
  • Final Dashboard
  • AI interpretation of analysis results, under the rule that AI must never invent numbers
Page 28 of 28

12. Glossary

Analis Data Retail — The persona who uploads, inspects, validates, and cleans the business dataset.

Pemilik/Pengelola Bisnis Retail — The persona who owns or manages the retail business and judges whether the data is trustworthy enough to use.

Data Workspace — The part of the application where business datasets are brought in, listed, inspected, and validated.

Dataset — A single uploaded business file (CSV or XLSX) together with its metadata, status, preview, and validation state.

Dataset type — The classification of a dataset under the data contract; in this delivery, the sales dataset.

Canonical fields — The agreed field names for a dataset type. For the sales dataset: tanggal, transaksi_id, pos, produk, kategori, brand, unit, harga, diskon, revenue, customer_id.

REQUIRED — A canonical field that must be present for the dataset to satisfy the contract.

OPTIONAL — A canonical field that may be absent without the dataset failing the contract.

DERIVED — A canonical field the system computes rather than reads, for example revenue computed from unit and harga. Derived columns carry a permanent DERIVED badge.

Data contract — The definition of dataset type, required columns, optional columns, data types, and validation rules that validation and cleaning are applied against.

Raw dataset — The dataset exactly as uploaded, before cleaning.

Cleaned dataset — The dataset produced by the cleaning pipeline, retained alongside the raw dataset.

Validation report — The report explaining which contract rules were applied and what the outcome was for each.

Validation status — The summary of a dataset's validation state against the data contract.

Data issue — A detected data problem, reported with a severity, an affected-row count, and a human sentence.

Transformation record — The stored record of what each important transformation did, to which column, and affecting how many values.

Cleaning pipeline — The Python/Pandas flow RAW DATA → VALIDATION → CLEANING → PROCESSED DATA.

DATA FIRST — The principle that the dataset is the sole source of every number the system shows.

Deterministic analytics — The principle that Python/Pandas computes numbers and statistics first, and that AI interpretation comes second and never invents numbers.

WIB — Waktu Indonesia Barat, the timezone used throughout the interface.

Perlu tindakan — The tangerine severity marker for a validation finding that requires the user's attention.

Bersih — The lime status marker for a valid or clean state.

No completed page designs yet.

Completed design pages will appear here when they are ready to preview.

Landing: Open workspace, read empty state
Upload CSV: Select CSV file
Upload CSV: Resubmit rejected file
Datasets: Confirm dataset listed
Dataset Metadata: Read column inventory and roles
Dataset Preview: Confirm actual rows
Dataset Status: Run validation
Validation Status: Read validation summary
Validation Report: 1. Read contract findings
Data Issues: 2. Read problem explanations
Data Cleaning: 3. Run cleaning pipeline
Data Issues: Read transformation record
Dataset Preview: Compare cleaned preview
Data Cleaning: 4. Re-run cleaning after fixes
Validation Report: Re-run validation
Upload XLSX: Select XLSX file
Upload XLSX: Resubmit rejected file
Landing: Open Datasets section

No completed page designs yet.

Completed design pages will appear here when they are ready to preview.

Landing: Open workspace, read empty state
Upload CSV: Select CSV file
Upload CSV: Resubmit rejected file
Datasets: Confirm dataset listed
Dataset Metadata: Read column inventory and roles
Dataset Preview: Confirm actual rows
Dataset Status: Run validation
Validation Status: Read validation summary
Validation Report: 1. Read contract findings
Data Issues: 2. Read problem explanations
Data Cleaning: 3. Run cleaning pipeline
Data Issues: Read transformation record
Dataset Preview: Compare cleaned preview
Data Cleaning: 4. Re-run cleaning after fixes
Validation Report: Re-run validation
Upload XLSX: Select XLSX file
Upload XLSX: Resubmit rejected file
Landing: Open Datasets section