JAYAPURA RETAIL INTELLIGENCE (project name: retail-jayapura) is a web application intended to become an AI Retail Intelligence system for electronics/gadget retail businesses in Jayapura and the surrounding Papua region. Its long-term purpose is to help those businesses understand penjualan, permintaan pasar, kebutuhan dan keluhan pelanggan, teknologi, inventory, keuangan dasar, laporan Excel, business intelligence, and AI business analysis.
This document covers Master Prompt 1 only, which is delivered in three sequential phases within a single body of work:
The governing principle across all three phases is:
DATA FIRST. DETERMINISTIC ANALYTICS FIRST. AI INTERPRETATION SECOND.
Python/Pandas computes numbers and statistics first. AI will later only interpret analysis results and help produce insight/recommendation. AI must never invent numbers. In this master prompt, AI is not activated at all.
The audience is the retail data analyst who prepares and cleans the shop's own exports, and the retail business owner/manager who needs certainty that any number the system will eventually show comes from a real dataset rather than a demo figure.
The current delivery is a first-party web workspace with no login and no authentication. It consists of a React + TypeScript + Vite frontend, a Python + FastAPI backend, and a Pandas/NumPy/openpyxl data layer. There is no AI module active, no business database, no API key, and no fabricated business content of any kind.
The accepted behavior in this delivery is:
Rp0 as business data, never "100 transaksi", never a fake chart, never a sample product, never a fake KPI.Rp for Rupiah, DD/MM/YYYY for dates, and WIB as the timezone — but shows no business numbers at all while no dataset is available.Actors in this delivery are two human personas (Analis Data Retail, Pemilik/Pengelola Bisnis Retail) and the application's own backend/analytics processes. There are no external providers, no outbound recipients, and no third-party surfaces.
Narrow exclusions for this delivery: no login, no authentication, no AI agent, no chatbot, no web scraping, no real market data, no dummy sales data, no dummy revenue, no dummy customer, no dummy inventory, no business database, no API key, no fake business intelligence, and no sample/random/hardcoded business numbers. Also excluded from this delivery and reserved for later master prompts: a complete Sales Analytics Engine, Market Intelligence, Customer Intelligence, Technology Intelligence, Inventory Analytics, Accounting, AI Business Analyst, Report Generator, and Final Dashboard.
This delivery is a data custody instrument, not an analytics product. Its whole promise is that the user can trust what the system says about their own files. That promise is why the workspace opens on an honest absence rather than a dashboard, why every count and size is a real measurement of a real file, and why validation findings are stated as sentences about the user's data rather than as decorative alerts.
Delivery ownership. All eleven surfaces in this delivery are first-party application pages. There is no provider-owned surface, no external destination, and no headless-only delivery. The backend and the Pandas cleaning pipeline are system processes that support the human-facing pages; they are never the only owner of a human-facing capability.
Access ownership. Access to the workspace requires no login and no authentication. This is an explicit source constraint, not an omission: the user forbade building login and authentication in this master prompt. Every page in this delivery is therefore reachable without establishing identity, and no page in this delivery owns an identity-establishment interaction. There is no application-owned identity, no session continuity requirement, and no differentiated permission or role-based visibility anywhere in this delivery. The two personas differ by what they need from the workspace, not by what they are permitted to see.
Current vs. future boundary. Current: foundation scaffold, Data Workspace (upload CSV, upload XLSX, listing, status, metadata, preview, validation status), data contract, validation and cleaning pipeline, validation report, transformation record, honest empty states, basic tests, README. Future (explicitly deferred, not built here): the Sales Analytics Engine, Market Intelligence, Customer Intelligence, Technology Intelligence, Inventory Analytics, Accounting, AI Business Analyst, Report Generator, and Final Dashboard — all to be done in later master prompts. AI interpretation is deferred. A business database is deferred unless the architecture genuinely requires one.
Stop condition. After Master Prompt 1 is complete, work stops. Master Prompt 2 is not started before the next instruction is given.
Not applicable. No reference directive in this request declares content_source authority, so no source content inventory is produced. All product facts in this document come from the authoritative user requirement thread and the accepted Planning Scope.
The page inventory below is the closed, ordered page contract for this delivery. Each page appears exactly once.
DERIVED acid-lime monospace chip beside any canonical field the system computed rather than read.Perlu tindakan, lime Bersih, muted Info); tabular affected-row counts.FR-1. Project scaffold structure — explicit
As a developer, I should have a project rooted at jayapura-retail-intelligence/ containing frontend/, backend/, analytics/, ai/, data/raw/, data/processed/, data/external/, reports/, tests/, docs/, README.md, .gitignore, and .env.example, so that the project has a clean and professional foundation.
FR-2. Frontend stack — explicit As a developer, I should have the frontend built with React, TypeScript, and Vite, so that the interface is developed on the specified stack.
FR-3. Backend stack — explicit As a developer, I should have the backend built with Python and FastAPI, so that the API is developed on the specified stack.
FR-4. Data stack — explicit As a developer, I should have the data layer built with Pandas, NumPy, and openpyxl, so that dataset reading, validation, and cleaning use the specified libraries.
FR-5. AI not activated — explicit As a developer, I should have no AI activated in this master prompt, so that the delivery stays within the DATA FIRST boundary.
FR-6. Database deferred — explicit As a developer, I should not introduce a database at this stage unless the architecture genuinely requires one, so that the delivery stays minimal.
FR-7. Extensible code structure — explicit As a developer, I should have a code structure that is easy to extend, so that later master prompts can build on it.
FR-8. Phase 1 exclusions — explicit As a developer, I should not build login, authentication, AI agent, chatbot, web scraping, real market data, dummy sales data, dummy revenue, dummy customer, dummy inventory, business database, API key, or fake business intelligence, so that the foundation contains no prohibited capability.
FR-9. CSV dataset upload — explicit As an Analis Data Retail, I should be able to upload a business dataset in CSV format, so that my own sales data can enter the workspace.
FR-10. XLSX dataset upload — explicit As an Analis Data Retail, I should be able to upload a business dataset in XLSX format, so that my own sales data can enter the workspace.
FR-11. File type validation on upload — explicit As an Analis Data Retail, I should have the system validate the file type before accepting a file as a business dataset, so that arbitrary files are not treated as business data.
FR-12. Clean and safe upload endpoints — explicit As a developer, I should provide clean and safe backend endpoints for the upload feature where the upload requires a backend API, so that ingestion is handled properly.
FR-13. Dataset listing — explicit As an Analis Data Retail, I should see a listing of the datasets available in the Data Workspace, so that I know what data I have brought in.
FR-14. Dataset status — explicit As an Analis Data Retail, I should see the processing and validation status of a dataset, so that I know where it stands in the pipeline.
FR-15. Dataset metadata — explicit As an Analis Data Retail, I should see metadata for a dataset, including its dataset type and data structure, so that I understand what I uploaded.
FR-16. Dataset preview — explicit As an Analis Data Retail, I should see a preview of the actual rows of a dataset, so that I can confirm the file contains what I expect.
FR-17. Data validation status — explicit As an Analis Data Retail, I should see the validation status of a dataset against the data contract, so that I know whether it is usable.
FR-18. Honest empty state when no dataset exists — explicit As a Pemilik/Pengelola Bisnis Retail, I should see an honest empty state stating that no dataset is available when no dataset has been uploaded, so that I am never shown fabricated business information.
Rp0 as business data, "100 transaksi", a fake chart, a sample product, or a fake KPI.FR-19. No sample data — explicit As a Pemilik/Pengelola Bisnis Retail, I should never see example data in the workspace, so that everything I see is my own.
FR-20. Cleaning pipeline flow — explicit As an Analis Data Retail, I should have the data flow RAW DATA → VALIDATION → CLEANING → PROCESSED DATA implemented with Python/Pandas, so that my raw file becomes a processed dataset through a defined sequence.
FR-21. No business analysis in Phase 3 — explicit As a developer, I should not perform business analysis in Phase 3, so that cleaning stays separate from analysis.
FR-22. Detection of missing required columns — explicit As an Analis Data Retail, I should have the system detect that a required column is missing, so that I know the dataset cannot satisfy the contract.
FR-23. Detection of mismatched column names — explicit As an Analis Data Retail, I should have the system detect column names that do not match the contract, so that I can correct my headers.
FR-24. Detection of invalid dates — explicit As an Analis Data Retail, I should have the system detect invalid date values, so that I know which rows cannot be read as dates.
FR-25. Detection of numbers read as text — explicit As an Analis Data Retail, I should have the system detect numeric values that were read as text, so that I can fix the source formatting.
FR-26. Detection of empty unit — explicit
As an Analis Data Retail, I should have the system detect empty unit values, so that I know where quantity is missing.
FR-27. Detection of empty harga — explicit
As an Analis Data Retail, I should have the system detect empty harga values, so that I know where price is missing.
FR-28. Detection of empty revenue — explicit
As an Analis Data Retail, I should have the system detect empty revenue values, so that I know where revenue is missing.
FR-29. Detection of duplicate transactions — explicit As an Analis Data Retail, I should have the system detect duplicate transactions, so that I know where the same transaction appears more than once.
FR-30. Detection of suspicious negative values — explicit As an Analis Data Retail, I should have the system detect suspicious negative values, so that I can review them before they affect anything downstream.
FR-31. Detection of differing date formats — explicit As an Analis Data Retail, I should have the system detect that date formats differ within a column, so that I know the column is not uniform.
FR-32. Detection of whitespace — explicit As an Analis Data Retail, I should have the system detect whitespace problems in values, so that inconsistent text does not silently split categories.
FR-33. Detection of inconsistent category names — explicit
As an Analis Data Retail, I should have the system detect inconsistent kategori names, so that the same category is not counted as several.
FR-34. Detection of inconsistent product names — explicit
As an Analis Data Retail, I should have the system detect inconsistent produk names, so that the same product is not counted as several.
FR-35. Raw dataset, cleaned dataset, and validation report concepts — explicit As an Analis Data Retail, I should have a raw dataset, a cleaned dataset, and a validation report as distinct concepts, so that I can always compare what I uploaded with what was produced.
FR-36. Explain data problems to the user — explicit As an Analis Data Retail, I should have the system explain data problems to me in plain language, so that I understand what is wrong with my file.
FR-37. No silent deletion — explicit As an Analis Data Retail, I should never have data deleted silently, so that I retain control over my own file.
FR-38. Record important transformations — explicit As an Analis Data Retail, I should have the system record what it did for every important transformation, so that the cleaned dataset is auditable.
FR-39. Data contract definition — explicit As a developer, I should have a clear data contract defining dataset type, required columns, optional columns, data types, and validation rules, so that validation and cleaning have a single authoritative definition.
FR-40. Sales dataset canonical fields — explicit
As a developer, I should define canonical fields for the sales dataset as tanggal, transaksi_id, pos, produk, kategori, brand, unit, harga, diskon, revenue, and customer_id, so that sales data has a stable vocabulary.
FR-41. Required, optional, and derived distinction — explicit As a developer, I should distinguish REQUIRED, OPTIONAL, and DERIVED fields, and not force every field to be present when it is not needed, so that the contract fits real files.
FR-42. Revenue as a derived field — explicit
As an Analis Data Retail, I should have revenue treated as DERIVED when unit and harga are available, so that revenue can be computed rather than required in my file.
revenue is absent but unit and harga are present, revenue is computed and the column is tagged DERIVED.FR-43. All numbers computed from the dataset — explicit As a Pemilik/Pengelola Bisnis Retail, I should have every number that comes from a dataset genuinely computed from that dataset, so that I can trust it.
FR-44. No demo numbers, random data, or hardcoded business numbers — explicit As a Pemilik/Pengelola Bisnis Retail, I should never see a number created for a demo, generated randomly, or hardcoded as business data, so that nothing in the workspace is fabricated.
FR-45. Indonesian interface language — explicit As an Analis Data Retail, I should have the main interface in Indonesian, so that I can work in the language of my business.
FR-46. Indonesian number and date formats and WIB timezone — explicit
As an Analis Data Retail, I should see Rupiah formatted with Rp, dates formatted as DD/MM/YYYY, and times in WIB, so that values match local convention.
Rp, dates use DD/MM/YYYY, and timezone is WIB.FR-47. No business numbers before a dataset exists — explicit As a Pemilik/Pengelola Bisnis Retail, I should see no business numbers at all while no dataset is available, so that the absence is stated rather than filled with a placeholder.
Rp0, no transaction count, no chart, no sample product, and no KPI is rendered before a real dataset exists.FR-48. Simple, professional, extensible UI — explicit As a developer, I should have a UI that is simple, professional, and easy to extend, so that later master prompts can build on it.
FR-49. File validation tests — explicit As a developer, I should have basic tests for file validation, so that rejected file types are proven to be rejected.
FR-50. Dataset validation tests — explicit As a developer, I should have basic tests for dataset validation, so that contract checks are proven.
FR-51. Column validation tests — explicit As a developer, I should have basic tests for column validation, so that missing and mismatched columns are proven to be detected.
FR-52. Data type validation tests — explicit As a developer, I should have basic tests for data type validation, so that type problems are proven to be detected.
FR-53. Cleaning pipeline tests — explicit As a developer, I should have basic tests for the cleaning pipeline, so that the RAW → VALIDATION → CLEANING → PROCESSED flow is proven.
FR-54. Empty dataset state tests — explicit As a developer, I should have basic tests for the empty dataset state, so that the honest empty state is proven.
FR-55. Backend runs, frontend builds, no serious lint errors — explicit As a developer, I should have the Python backend runnable, the frontend buildable, and no serious lint errors, so that the project is in a working state.
FR-56. Honest test reporting — explicit As a developer, I should report test results honestly without inventing them, and fix failures where possible, so that the reported state matches reality.
FR-57. README content — explicit As a developer, I should have the README updated with the project purpose, architecture, folder structure, data flow, data contract, roadmap, the DATA FIRST principle, the deterministic analytics principle, and the status of completed phases, so that the project is documented.
FR-58. Master Prompt 1 final report — explicit As a developer, I should deliver the final structure, the list of important files created, the architecture explanation, the Data Workspace explanation, the cleaning pipeline explanation, the data contract explanation, the list of endpoints created, the list of tests, the lint/build/test results, and the list of things NOT yet built, so that the work is fully accounted for.
FR-59. Stop after Master Prompt 1 — explicit As a developer, I should stop after Master Prompt 1 and not begin Master Prompt 2 before the next instruction, so that the delivery boundary is respected.
Product context. This persona prepares the shop's own data. They export sales records from a POS or spreadsheet, and the file they hold is rarely clean: headers drift between exports, dates arrive in more than one format, quantities are blank, and the same product is spelled three ways. They are the person who has to make that file trustworthy before anyone else looks at it.
Primary goal. Bring a CSV or XLSX business dataset into the workspace, understand exactly what is wrong with it, and produce a cleaned dataset with a validation report and a record of what was changed — without any business analysis being performed yet.
Distinct accepted responsibilities.
Relevant inputs and decisions. The input is the user's own file. The decisions are: whether the file is the right one (confirmed by preview), whether the detected problems are acceptable or must be fixed at source, and whether to proceed to cleaning. The persona decides when a dataset is good enough to move forward.
Interactions with other accepted participants. This persona produces the cleaned dataset and the validation report that the Pemilik/Pengelola Bisnis Retail reads to judge whether the data is fit to use. The analyst's work is the precondition for the owner's confidence.
Observable success. A cleaned dataset exists, a validation report explains the problems found, and a transformation record shows what was done — with every count traceable to the uploaded file.
What makes this role different. This persona operates on the data itself: they upload, inspect, validate, and clean. Their work is the pipeline.
Product context. This persona owns or manages an electronics/gadget retail business in Jayapura or the surrounding area. They are not going to clean files themselves, but they are the person who will eventually act on numbers, and they have been burned before by dashboards that showed impressive figures that turned out to be examples.
Primary goal. Be certain that any number the system will eventually show comes from a real dataset rather than a demo, and that the system says plainly when there is no data.
Distinct accepted responsibilities.
Relevant inputs and decisions. The input is the analyst's uploaded dataset and its validation outcome. The decision is whether the data is trustworthy enough to be used as the basis for later analysis.
Interactions with other accepted participants. This persona depends on the Analis Data Retail's upload and cleaning work. They consume the validation report and the data issues list; they do not produce them.
Observable success. The workspace shows either a real dataset with real measured counts, or a plain statement that no dataset is available. At no point does it show Rp0 as business data, a transaction count, a chart, a sample product, or a KPI that was not computed from a real file.
What makes this role different. This persona does not operate on the data; they audit the system's honesty about the data. Their success condition is the absence of fabrication, which is a different kind of requirement from the analyst's success condition of a clean dataset.
Rp0, no transaction count, no chart, no sample product, and no KPI is rendered.revenue was absent but unit and harga were present, the revenue column carries the acid-lime DERIVED chip.Perlu tindakan, lime Bersih, or muted Info, and affected-row counts are set in tabular numerals.unit, empty harga, empty revenue, duplicate transactions, suspicious negative values, differing date formats, whitespace, inconsistent kategori names, and inconsistent produk names.Rp0 presented as business data, no "100 transaksi", no fake chart, no sample product, and no fake KPI.DERIVED chip.Muse and headline. Rasmus Andersson — systematic product craft with an opinion. The register is custody, not excitement: this tool must read as a precision instrument that will not lie about the user's data.
Mode. Dark mode only.
Colour tokens by role.
| Role | Token | Value |
|---|---|---|
| Page ground | --bg | #101113 |
| Panel / table surface | --surface | #16181B |
| Ink (primary text) | --text | #EDEAE4 |
| Signal (primary) | --primary | #FF6B2C |
| Valid / derived status | --accent | #C6F24E |
| Muted (labels, metadata keys, timestamps, empty-state copy) | --muted | #8A8A85 |
| Hairline border | --hairline | rgba(237,234,228,0.10) |
#101113 is the page ground — warm graphite, never pure black. #16181B is the panel and table surface, separated from the ground by 1px hairlines. #EDEAE4 is warm off-white ink at 14.5:1 on the ground. #FF6B2C is THE signal: primary buttons, focus rings, the active upload dropzone border, the current nav underline, and the single Perlu tindakan severity marker — used on roughly 3% of pixels. #C6F24E is a rarer second status colour reserved exclusively for valid / bersih state chips and the DERIVED field badge, never decorative. #8A8A85 carries labels, metadata keys, timestamps, and honest empty-state copy.
Ratio. Approximately 78% graphite ground, 15% panel surface, 4% ink, 3% tangerine and lime combined.
No blue. There is no blue, indigo, or violet anywhere in the system, including links. Links are ink with a tangerine underline.
Typography.
clamp), section head 24/28, panel title 18, body 15, table cell 13.5, micro-label 11 uppercase.tabular-nums.Shape language. Rectilinear and honest. 6px radius on controls, 8px on panels, 3px on chips. No pill buttons, no blobs, no soft offset shadows. Separation comes from 1px hairlines and one-step surface lifts, never from elevation blur. The upload dropzone is a dashed 1px rectangle, becoming 2px dashed tangerine only while dragging over it. Table rows are ruled, not carded; row hover is a flat surface fill, not a lift.
Spacing rhythm. 4/8-pt grid throughout. Main column is a strict 12-column grid, max-width 1440px, 24px gutters mobile / 32px desktop.
Layout. A fixed left rail — 72px icon rail at 375–767px, 232px labelled rail at ≥1024px — carrying the workspace sections: Ringkasan Data, Unggah Dataset, Daftar Dataset, Validasi & Pembersihan, Kontrak Data. The dataset detail view is a two-pane split: metadata and validation report on the left (5 columns, sticky), raw/cleaned preview table on the right (7 columns, horizontally scrollable with a frozen first column). Every panel header is a ruled bar with the title left and a status chip right. Nothing is centred; everything is flush-left on the grid.
Imagery style. The interface is the imagery. No photography, no illustration, no 3D. Visual interest is carried by tabular numeral blocks, a monospace schema listing of the data contract, hairline-grid diagrams showing the four-stage pipeline as ruled boxes with arrows, and a sparkline-free distribution strip of column fill rates rendered as a row of 1px vertical bars. The empty state is a single centred 1px-ruled rectangle with muted copy — an absent shelf, not a cartoon.
Accessibility. Text and controls stay whole at 375px, 768px, and 1280px: headlines, labels, numbers, and controls remain entirely inside the viewport and their container, wrapping or scaling to fit, and no other element covers any part of them. Decoration may be cropped or bled off an edge; readable text and controls may not.
The honest absence. The first screen is the workspace itself, opened on Ringkasan Data with zero datasets loaded. There is no marketing hero.
The composition: a full-height graphite field (#101113), a 72px icon rail hard against the left edge, and in the main column a flush-left display line in Space Grotesk at 40px mobile → 72px desktop reading "Belum ada dataset." set on two lines, the second line in muted #8A8A85. Directly beneath it, a single 1px-ruled horizontal rule spanning the full content width. Beneath that, one sentence in body copy: "Unggah file CSV atau XLSX untuk memulai. Semua angka akan dihitung dari data Anda, bukan dari contoh." Beneath that, one solid tangerine rectangle button (Unggah Dataset) 44px tall, flush-left, and to its right a dashed 1px outlined secondary target.
The right two-thirds of the viewport stays empty graphite. That negative space is deliberate: it reads as an instrument waiting for a specimen. No KPI tiles, no Rp0, no charts, no illustration, no gradient.
This concept recomposes only accepted content and controls — the honest empty state, the upload entry, and the workspace shell. It introduces no new behavior, page, or destination.
Interaction Model: Static Motion Tempo: restrained Hero Dimensionality: flat
Landing Hero Motion Brief
cubic-bezier(0.2, 0, 0, 1). Upload progress is a determinate tangerine bar, not a spinner. Validation report rows stagger in at 40ms intervals on first paint and then never animate again. Table row hover is an instant surface fill (90ms). Focus rings appear in 0ms. No parallax, no scroll reveals, no bounce, no counters ticking up.prefers-reduced-motion, every transition collapses to 0ms and the validation-report stagger becomes a single paint. The first frame is unchanged.No 3D or WebGL scene is required or requested for this delivery.
NFR-1. Deterministic computation before interpretation — explicit Python/Pandas must compute numbers and statistics first; AI interpretation comes second and must never invent numbers. In this delivery AI is not activated at all. Rationale: this is the governing product principle.
NFR-2. No fabricated business content — explicit No demo numbers, no random data, and no hardcoded business figures may exist anywhere in the system. Rationale: the product's entire value is that the user can trust what it says about their own data.
NFR-3. Honest empty state — explicit
When no dataset exists, the system must state that the dataset is not yet available. It must not render Rp0 as business data, "100 transaksi", a fake chart, a sample product, or a fake KPI. Rationale: an empty state that lies destroys the custody promise.
NFR-4. File type validation — explicit Uploads must validate file type; arbitrary files must not be accepted as business datasets. Rationale: the workspace must not ingest non-business files.
NFR-5. No silent data deletion — explicit Data must never be removed silently; problems must be explained to the user. Rationale: the user retains custody of their file.
NFR-6. Transformation traceability — explicit Every important transformation must be recorded with what was done. Rationale: the cleaned dataset must be auditable against the raw dataset.
NFR-7. No forced fields — explicit The data contract must not force every field to be present when it is not needed. Rationale: real exports vary, and a contract that rejects valid files is useless.
NFR-8. Honest test reporting — explicit Test results must be reported honestly and never invented; failures must be disclosed and fixed where possible. Rationale: the project's credibility depends on truthful reporting.
NFR-9. Indonesian locale conventions — explicit
The main interface is in Indonesian; currency uses Rp; dates use DD/MM/YYYY; timezone is WIB. Rationale: the audience operates in Indonesian retail convention.
NFR-10. Runnable backend, buildable frontend, clean lint — explicit The Python backend must run, the frontend must build, and there must be no serious lint errors. Rationale: the project must be in a working state.
NFR-11. No login or authentication — explicit No login and no authentication may be built in this delivery. Rationale: explicitly excluded by the user for this master prompt.
NFR-12. No AI, no database, no API key — explicit AI is not activated; a database is not required unless the architecture genuinely requires one; no API key is created. Rationale: explicitly excluded or deferred by the user for this master prompt.
NFR-13. Readable text and controls at every viewport — required_inference Headlines, labels, numbers, and controls must remain entirely inside the viewport and their container at 375px, 768px, and 1280px, wrapping or scaling to fit, with no other element covering them. Rationale: required to make the accepted interface usable at the stated breakpoints.
All technology choices below are explicit user requirements.
Frontend
Backend
Data
AI
Database
Project structure
jayapura-retail-intelligence/
├── frontend/
├── backend/
├── analytics/
├── ai/
├── data/
│ ├── raw/
│ ├── processed/
│ \x20\xe2\x94\x94── external/
├── reports/
├── tests/
├── docs/
├── README.md
├── .gitignore
\xe2\x94\x94── .env.example
Containerization and orchestration
[Default — not specified by user] Docker and docker-compose for local development and running the frontend, backend, and analytics together.[Default — not specified by user] Kubernetes is not required for this delivery.Assumptions
[Assumption] The user's CSV and XLSX files are exports from their own POS or spreadsheet tooling and are not encrypted or password-protected.[Assumption] A single uploaded file corresponds to a single dataset.[Assumption] The sales dataset is the only dataset type defined in this delivery; the data contract mechanism is general enough to add further dataset types later.[Assumption] The workspace is used by a small number of people in the same business, which is why no login and no differentiated permissions are needed in this delivery.[Assumption] The ai/ directory exists as structure only and contains no active AI behavior.Constraints
Rp0 as business data, "100 transaksi", a fake chart, a sample product, or a fake KPI when no dataset is available.Future requirements (not current, not in scope for this delivery)
Analis Data Retail — The persona who uploads, inspects, validates, and cleans the business dataset.
Pemilik/Pengelola Bisnis Retail — The persona who owns or manages the retail business and judges whether the data is trustworthy enough to use.
Data Workspace — The part of the application where business datasets are brought in, listed, inspected, and validated.
Dataset — A single uploaded business file (CSV or XLSX) together with its metadata, status, preview, and validation state.
Dataset type — The classification of a dataset under the data contract; in this delivery, the sales dataset.
Canonical fields — The agreed field names for a dataset type. For the sales dataset: tanggal, transaksi_id, pos, produk, kategori, brand, unit, harga, diskon, revenue, customer_id.
REQUIRED — A canonical field that must be present for the dataset to satisfy the contract.
OPTIONAL — A canonical field that may be absent without the dataset failing the contract.
DERIVED — A canonical field the system computes rather than reads, for example revenue computed from unit and harga. Derived columns carry a permanent DERIVED badge.
Data contract — The definition of dataset type, required columns, optional columns, data types, and validation rules that validation and cleaning are applied against.
Raw dataset — The dataset exactly as uploaded, before cleaning.
Cleaned dataset — The dataset produced by the cleaning pipeline, retained alongside the raw dataset.
Validation report — The report explaining which contract rules were applied and what the outcome was for each.
Validation status — The summary of a dataset's validation state against the data contract.
Data issue — A detected data problem, reported with a severity, an affected-row count, and a human sentence.
Transformation record — The stored record of what each important transformation did, to which column, and affecting how many values.
Cleaning pipeline — The Python/Pandas flow RAW DATA → VALIDATION → CLEANING → PROCESSED DATA.
DATA FIRST — The principle that the dataset is the sole source of every number the system shows.
Deterministic analytics — The principle that Python/Pandas computes numbers and statistics first, and that AI interpretation comes second and never invents numbers.
WIB — Waktu Indonesia Barat, the timezone used throughout the interface.
Perlu tindakan — The tangerine severity marker for a validation finding that requires the user's attention.
Bersih — The lime status marker for a valid or clean state.
No completed page designs yet.
Completed design pages will appear here when they are ready to preview.
No completed page designs yet.
Completed design pages will appear here when they are ready to preview.
No comments yet. Be the first!