Page 1 of 15
System Requirements Document
1. Introduction
This System Requirements Document (SRD) specifies the requirements for an empirical corporate-disclosure research system: a replicable, source-registered panel study that investigates whether the public availability of financial and non-financial corporate disclosure is associated with subsequent stock-price outcomes and trading activity among non-financial issuers listed on the Egyptian Exchange (EGX), together with the scholarly editorial production of the resulting manuscript.
The system produces two coupled deliverables:
- A quantitative, longitudinal, associational research pipeline that constructs a firm-year and firm-day panel spanning 2016–2025, codes two archive-availability disclosure indices from public issuer and investor-relations archives, estimates pooled benchmark regressions and two-way fixed-effects models across four market outcomes, and subjects the estimates to a pre-specified inference and robustness battery.
- A manuscript editorial workflow that elevates the resulting academic text to the linguistic, structural, and evidentiary standards of elite high-impact business and management journals, including systematic exhibit verification and an original-versus-revised edit presentation. The workflow is executed under the persona of a distinguished finance and business scholar, senior editor, and peer reviewer with over 30 years of continuous publication experience in the top 5% of high-impact business and management journals (e.g., Academy of Management Journal, Strategic Management Journal, Administrative Science Quarterly), and its authoritative content source is the uploaded revised manuscript (revised_manuscript_Financial_and_Non-Financial_Disclosure_4.docx), which is the text to be edited.
The system is explicitly associational, not causal. All estimates are conditional associations. Archive-availability indices record whether a disclosure item was located in cited public archives; they do not measure disclosure quality, timeliness, or content.
Page 2 of 15
2. System Overview
The system is a reproducible empirical-research stack rather than an interactive software product. Its operational shape is:
- Ingestion layer — Yahoo Finance daily equity pricing series (close, adjusted close, volume) for eligible EGX issuers, and World Bank Development Indicators macro series (inflation, official EGP/USD rate, GDP growth), plus the EGX30 annual return series.
- Coding layer — a manual archive-availability coding protocol applied to public issuer filings and investor-relations archives, producing a four-item Financial Disclosure Index (FDI) and a three-item Non-Financial Disclosure Index (NFDI), lagged one year, with row-level source URLs and retrieval dates.
- Computation layer — pooled benchmark regressions (pooled OLS), two-way (firm and year) fixed-effects panel models, and a daily firm-date extension estimated in Python, with firm-clustered standard errors, small-cluster bootstrap inference, and multiple-testing correction.
- Validity layer — pre-estimation diagnostics (unit root, multicollinearity, panel specification) and post-estimation robustness (placebo leads, winsorization, subsample, leave-one-sector-out, correlated-random-effects cross-check).
- Reporting layer — descriptive statistics, correlation and VIF exhibits, regression tables, coefficient plots, a sample-selection waterfall, a limitations catalogue, and a recommendation set bounded by the associational design.
- Editorial layer — a manuscript revision workflow applying sophisticated academic voice, structural framing, exhibit verification, and argumentative rigor, delivered as original-versus-revised text with editorial notes. The layer operates on the authoritative manuscript source, preserves the manuscript's substantive content and meaning, and produces an auditable edit record in which every edit is paired with its rationale.
- Replication layer — deposited datasets, data dictionary, source register, coding log, and replication code with an environment lockfile and README.
The system intentionally contains no interactive front end, user authentication, administration surface, or product account model. Its outputs are static analytical artifacts, exhibits, and manuscript text.
3. Functional Requirements
Page 3 of 15
3.1 Population, Sample, and Panel Construction
- As a researcher, I want to define the target population as non-financial common-equity operating companies listed on the EGX at any point during 2016–2025, so that the study window spans a full macroeconomic crisis cycle.
- As a researcher, I want to exclude banks, insurance companies, brokerage houses, investment funds, and leasing entities, so that issuers with distinct regulatory disclosure regimes do not contaminate the sample.
- As a researcher, I want to require an identifiable ticker with a Yahoo Finance history, so that daily market outcomes can be computed.
- As a researcher, I want to require at least five calendar years with ≥20 valid trading observations each, so that firms with thin or interrupted trading histories are excluded.
- As a researcher, I want to require non-missing close, adjusted close, and volume fields, so that return, volatility, volume, and turnover measures are computable.
- As a researcher, I want to require an available public disclosure archive, so that archive-availability coding can be performed.
- As a researcher, I want to retain the complete universe of 34 eligible non-financial EGX issuers with surviving, continuous public archive data, so that the sample is framed as exhaustive rather than as a random subset.
- As a researcher, I want a sample-selection waterfall (Figure 1) tracing the inclusion/exclusion path, so that stage-by-stage exclusion counts are transparent.
- As a researcher, I want to construct an unbalanced panel of 34 firms and 322 firm-year observations (78,037 daily firm-date observations), so that both annual and daily estimation are supported.
- As a researcher, I want calendar-year lagging of the disclosure indices to yield 288 lagged firm-years across 34 firms, so that temporal ordering of disclosure before outcomes is enforced.
- As a researcher, I want a pre-specified common-sample rule that excludes the four ORAS firm-years (2022–2025) with zero annual trading volume (for which log volume is undefined), yielding 284 observations across 33 firms for the headline results.
- As a researcher, I want an alternative outcome-specific rule that retains 288 observations for the return and volatility equations, so that both sample-rule results are reported in each regression table.
- As a researcher, I want data quality checks confirming zero duplicate firm-date keys, zero blank cells, and zero formula errors, so that the panel is internally consistent.
- As a researcher, I want survivorship and archive-availability selection to be acknowledged as sampling limitations rather than imputed away, so that the external-validity boundary is explicit.
3.2 Disclosure Index Construction and Coding
- As a researcher, I want to construct a four-item Financial Disclosure Index (FDI) = (audited statements + annual report + earnings announcement + ratios/EPS)/4, so that the financial disclosure construct is operationalized as archive availability.
- As a researcher, I want to construct a three-item Non-Financial Disclosure Index (NFDI) = (ESG/CSR + corporate governance + board/stakeholder)/3, so that the voluntary disclosure construct is operationalized on the same basis.
- As a researcher, I want to code both indices from public issuer and investor-relations archives following the index-construction tradition of Botosan (1997) and Clarkson et al. (2008), so that coding is comparable to prior disclosure-index literature.
- As a researcher, I want the indices to record availability in cited public archives rather than verified content quality or timeliness, so that the measurement construct is unambiguous.
- As a researcher, I want a coding value of zero to denote "not located in the cited public archives" rather than verified non-disclosure, so that the measure is not over-interpreted.
- As a researcher, I want to retain row-level source URLs and retrieval dates, so that every coded cell is traceable to its archival source.
- As a researcher, I want to lag both disclosure indices one year relative to the outcome period, so that simultaneity between disclosure choices and contemporaneous performance is addressed.
- As a researcher, I want to document all variables, sources, cleaning rules, and transformations row-level in a data dictionary and source register, so that the dataset is fully reconstructable.
Page 4 of 15
3.3 Market Outcome and Control Variable Computation
- As a researcher, I want to compute the annual adjusted return as the percentage change between the first and last valid adjusted closes per firm-year, so that the primary outcome is defined consistently.
- As a researcher, I want to compute annualized daily return volatility as the within-year standard deviation of log returns multiplied by √252, so that the risk outcome is annualized on a common basis.
- As a researcher, I want to compute log annual trading volume, so that the trading-activity outcome is estimated on a scale-stable basis.
- As a researcher, I want to compute annual turnover as annual volume divided by the public share-count denominator, so that a relative activity measure complements absolute volume.
- As a researcher, I want to include the firm-varying log market-capitalization proxy as a control, so that firm size is held constant within the estimated models.
- As a researcher, I want to include the EGX30 annual return (StockQ, 2026) as a macro control in pooled models only, so that market-wide return variation is conditioned upon where identified.
- As a researcher, I want to include World Bank inflation (FP.CPI.TOTL.ZG), the official EGP/USD rate (PA.NUS.FCRF), and GDP growth (NY.GDP.MKTP.KD.ZG) as macro controls in pooled models only, so that the macroeconomic regime is controlled where identification permits.
- As a researcher, I want to include industry indicators in pooled models and let firm fixed effects absorb industry classification in panel models, so that the design is internally coherent across specifications.
3.4 Econometric Estimation
- As a researcher, I want pooled benchmark regressions estimated across all four market outcomes, so that the cross-sectional disclosure–outcome associations are documented as a comparison baseline before within-firm identification is introduced.
- As a researcher, I want the pooled benchmark regressions specified as pooled OLS with industry indicators, year indicators, macroeconomic controls, and firm-clustered standard errors, so that the pooled design is estimated under a fully declared specification.
- As a researcher, I want the pooled benchmark regressions reported in a dedicated regression table covering all four outcomes (return, volatility, log volume, turnover), so that the benchmark estimates are directly comparable across outcomes.
- As a researcher, I want the pooled benchmark regression table to report coefficients, firm-clustered standard errors, adjusted R², F-statistics, and observation counts, so that the benchmark is assessable on standard reporting dimensions.
- As a researcher, I want the pooled benchmark regressions interpreted as vulnerable to omitted-variable bias from unobserved firm-level heterogeneity, so that they motivate rather than substitute for the fixed-effects framework.
- As a researcher, I want a two-way (firm and year) fixed-effects specification with firm-clustered standard errors as the preferred model, so that time-invariant firm heterogeneity and common macroeconomic shocks are absorbed.
- As a researcher, I want each of the four outcomes (annual adjusted return, annualized volatility, log annual volume, annual turnover) estimated separately in the fixed-effects framework, so that channel-specific associations are reported.
- As a researcher, I want annual macro variables excluded from the fixed-effects specification because they are common to all firms within a year and absorbed by year effects, so that the model does not purport to separately identify them.
- As a researcher, I want the pooled benchmark regressions and fixed-effects designs presented as alternatives rather than simultaneously identified specifications, so that estimation claims remain honest.
- As a researcher, I want a daily two-way fixed-effects extension estimated on 70,955 usable firm-days with annual indices carried forward to the firm-day level, so that a higher-frequency timing structure is examined and explicitly flagged as adding no independent identifying information.
Page 5 of 15
3.5 Model-Selection and Validity Diagnostics
- As a researcher, I want firm-level augmented Dickey–Fuller (ADF) tests on daily log returns constituting an Im–Pesaran–Shin-style panel unit-root screen, so that stationarity is established before inference.
- As a researcher, I want variance inflation factors (VIF) computed for all regression variables, so that multicollinearity is screened against conventional thresholds.
- As a researcher, I want Breusch–Pagan Lagrange Multiplier tests performed against random effects, so that the presence of unobserved firm heterogeneity is established.
- As a researcher, I want Hausman specification tests performed across all four outcome equations, so that fixed versus random effects is justified per outcome.
- As a researcher, I want Wooldridge tests for panel serial correlation, so that autoregressive dynamics in the outcome are detected.
- As a researcher, I want modified Wald tests for groupwise homoscedasticity, so that cluster-robust inference is motivated.
- As a researcher, I want a correlated-random-effects (Mundlak) specification reported in the robustness battery, so that the fixed-effects choice is cross-checked.
- As a researcher, I want the negative and non-positive-definite Hausman statistic for the log volume equation treated as uninformative, with two-way fixed effects retained throughout for design coherence, so that the diagnostics are reported transparently rather than suppressed.
3.6 Small-Cluster Inference and Multiple-Testing Control
- As a researcher, I want Rademacher wild-cluster bootstrap p-values (999 replications) reported alongside asymptotic cluster-robust p-values, so that inference with 33 firm clusters is not over-reliant on asymptotics.
- As a researcher, I want Benjamini–Hochberg adjusted p-values reported across the pre-specified family of eight hypothesis tests (two indices × four outcomes), so that multiple-testing inflation is controlled.
- As a researcher, I want annual adjusted return designated the primary outcome and volume, turnover, and volatility designated secondary tests, so that the testing family is pre-specified rather than data-dependent.
- As a researcher, I want nominal versus adjusted significance labeled throughout, so that readers cannot mistake nominal significance for confirmatory evidence.
3.7 Robustness and Sensitivity Battery
- As a researcher, I want disclosure-lead placebo tests for anticipation, so that one-period-ahead disclosure-to-outcome associations are ruled out for the primary outcome.
- As a researcher, I want the significant FDI lead in the volatility equation reported explicitly as evidence that the indices embed persistent firm risk traits, so that causal language is pre-emptively constrained.
- As a researcher, I want 1%/99% winsorization of annual returns applied to suppress extreme devaluation spikes, so that the primary association is shown not to be driven by tail observations.
- As a researcher, I want a hand-reviewed, document-verified subsample estimated separately (n = 36 firm-years, four firms), so that classical measurement error attenuation in the archive screen can be directionally assessed.
- As a researcher, I want the verified-subsample result reported without inference beyond the attenuation-bias interpretation, so that a four-firm result is not over-claimed.
- As a researcher, I want leave-one-sector-out re-estimation performed, so that sector-driven estimates are ruled out.
- As a researcher, I want a prespecified double-coding protocol (Cohen's κ for binary items, ICC for indices) specified for the verified subsample, with outstanding statistics reported as a limitation, so that content-analysis reliability standards are honored.
- As a researcher, I want both the common-sample and outcome-specific Ns reported in each regression table, so that sample-composition sensitivity is visible.
Page 6 of 15
3.8 Pre-Specified Hypotheses
- As a researcher, I want hypothesis H1 stating that lagged financial disclosure (FDI) is positively associated with subsequent annual stock returns, tested under the preferred specification.
- As a researcher, I want hypothesis H2 stating that lagged financial disclosure (FDI) is positively associated with subsequent trading volume and turnover, tested under the preferred specification.
- As a researcher, I want hypothesis H3 stating that lagged financial disclosure (FDI) is negatively associated with subsequent return volatility, tested under the preferred specification.
- As a researcher, I want hypothesis H4 stating that lagged non-financial disclosure (NFDI) is positively associated with subsequent annual stock returns, tested under the preferred specification.
- As a researcher, I want hypothesis H5 stating that lagged non-financial disclosure (NFDI) is positively associated with subsequent trading volume and turnover, tested under the preferred specification.
- As a researcher, I want hypothesis H6 stating that lagged non-financial disclosure (NFDI) is negatively associated with subsequent return volatility, tested under the preferred specification.
- As a researcher, I want the comparative-static prediction that the non-financial index (H4–H6) carries the stronger theoretical prior relative to the financial index (H1–H3) embedded in the design, so that the saturation asymmetry is a testable prediction rather than a post hoc observation.
3.9 Result Reporting and Interpretive Discipline
- As a researcher, I want the lagged NFDI–return association reported as suggestive rather than confirmatory, because it does not survive Benjamini–Hochberg adjustment.
- As a researcher, I want the economic magnitude of the NFDI–return association translated into a one-standard-deviation interpretation, so that the estimate is readable in economic terms.
- As a researcher, I want the null FDI–return association reported with its competing explanations (exhausted marginal information content versus insufficient remaining identifying variation) left undiscriminated, so that the data are not over-read.
- As a researcher, I want the marginal and adjustment-fragile NFDI–volume association reported as weak support, so that trading-activity claims remain bounded.
- As a researcher, I want the unsupported volatility hypotheses reported alongside the unconditional descriptive pattern (Figure 4), so that the dissipation of the unconditional pattern under fixed effects is transparent.
- As a researcher, I want a coefficient plot (Figure 5) consolidating two-way fixed-effects point estimates and firm-clustered 95% confidence intervals across all four outcomes, so that the result structure is visually summarized.
- As a researcher, I want a limitations, diagnostics, and residual-risks exhibit (Table 8) cataloguing each threat to inference, its likely bias direction, the mitigation applied, and the residual risk, so that boundary conditions are explicit.
- As a researcher, I want all conclusions framed as conditional associations rather than causal effects, so that the associational design does not yield causal language.
Page 7 of 15
3.10 Recommendations Output
- As a researcher, I want recommendations directed to the Financial Regulatory Authority and the EGX on phased sustainability-reporting deliberations aligned with ISSB standards (IFRS Foundation, 2023), explicitly stated as not justifying a regulatory mandate on associational evidence alone.
- As a researcher, I want an independent recommendation that machine-readable, timestamped disclosure registries be developed to improve market surveillance and enable future event-study research.
- As a researcher, I want recommendations for professional associations (auditors' and governance institutes) to develop sector-specific disclosure checklists covering environmental, employee, human-rights, board, and stakeholder items to convert availability into quality.
- As a researcher, I want recommendations for top management and boards on the capital-market relevance of non-financial disclosure beyond its reputational function and on documentable disclosure mattering more than archive presence.
- As a researcher, I want recommendations for institutional and individual investors noting that usability as a screening signal requires out-of-sample validation and content verification not performed by this study.
- As a researcher, I want a future-research agenda covering an exact-date disclosure event registry with market-model CARs and abnormal volume (MacKinlay, 1997), difference-in-differences or event-time designs around documented EGX/FRA regulatory changes, double-coded content indices reporting Cohen's κ and ICC plus NLP scoring of quality/tone/forward-looking content, reconstruction of historical issued and free-float shares, and extension to MENA cross-country panels testing whether institutional quality moderates the non-financial disclosure association.
3.11 Manuscript Editorial Production
- As a scholarly editor, I want to elevate the linguistic formulation of the manuscript to a sober, authoritative, and highly critical academic tone, so that the text meets the standards of the top 5% of high-impact business and management journals.
- As a scholarly editor, I want precise terminology, clear theoretical articulation, and robust empirical reporting throughout, so that conceptual and statistical claims are unambiguous.
- As a scholarly editor, I want redundancies eliminated, excessive passive voice reduced, and colloquialisms removed, so that the prose is tightened without altering meaning.
- As a scholarly editor, I want every section and sub-section to begin with a compelling opening sentence that frames the core argument, so that structural framing is consistent throughout.
- As a scholarly editor, I want every section and sub-section to conclude with a definitive closing sentence that synthesizes findings and provides a logical bridge to the next section, so that section boundaries carry argumentative momentum.
- As a scholarly editor, I want the transitions and section boundaries scrutinized as a distinct editorial pass, so that the framing and synthesis sentences are not merely present but argumentatively load-bearing at every boundary.
- As a scholarly editor, I want to systematically verify that every figure and table is explicitly, accurately, and organically cited within the narrative, so that no exhibit is orphaned.
- As a scholarly editor, I want professionally phrased citations inserted where an exhibit is implied but unreferenced (e.g., "As illustrated in Table 1..."), so that exhibit references are complete.
- As a scholarly editor, I want the logical flow of scientific arguments, theoretical models, concepts, and empirical results enhanced, so that the reasoning chain is coherent.
- As a scholarly editor, I want statistical findings discussed with precision and transparency, so that significance, effect size, and inference caveats are accurately conveyed.
- As a scholarly editor, I want the output presented as original text versus revised text for each edit, so that changes are auditable.
- As a scholarly editor, I want a brief "Editorial Notes" section below each revised passage detailing any specific structural adjustments made (such as added transitional sentences or corrected exhibit citations), so that the rationale for each edit is recorded.
- As a scholarly editor, I want the editorial workflow to operate on the authoritative uploaded manuscript as its content source, so that the revision is applied to the actual text under review rather than to a reconstructed or paraphrased substitute.
- As a scholarly editor, I want the editorial revision to preserve the manuscript's substantive content, empirical results, and meaning, so that the elevation of voice and structure does not alter findings, numbers, or claims.
Page 8 of 15
3.12 Declarations, Compliance, and Replication Deposits
- As a researcher, I want the derived annual and daily datasets, data dictionary, source register, coding log, and replication code (dataset build, estimation, and figure scripts with environment lockfile and README) deposited in a repository with a DOI, so that the study is replicable.
- As a researcher, I want raw issuer reports not redistributed where copyright restricts redistribution, with stable source URLs, publication dates, and retrieval dates provided instead, so that copyright constraints are respected.
- As a researcher, I want Yahoo Finance data treated as subject to the provider's terms, so that redistribution limits are observed.
- As a researcher, I want an ethics declaration stating that the study uses public secondary market and corporate-report data with no human participants or private identifiable data, so that ethics requirements are addressed.
- As a researcher, I want competing-interest, funding, and data-and-code-availability declarations included in the deliverable, so that journal disclosure obligations are met.
- As a researcher, I want every reference to carry a functioning DOI with all bracketed publication details resolved before submission, so that journal reference compliance is satisfied.
- As a researcher, I want the 2024–2026 Q1/Q2 literature base expanded to approximately twenty studies, so that the reference set meets reviewer recommendations.
4. User Personas
Primary personas (active, materially distinct workflows):
-
Empirical Disclosure Researcher (Principal Investigator / Author). Builds the EGX panel, applies the sample rules, codes the FDI and NFDI indices from issuer and investor-relations archives, computes market and macro variables, runs the pooled benchmark regressions and the two-way fixed-effects models, executes the diagnostics and robustness battery, and drafts the manuscript. Interacts with the ingestion, coding, computation, validity, and reporting layers.
-
Scholarly Editor / Peer Reviewer. A distinguished finance and business scholar, senior editor, and peer reviewer with over 30 years of continuous publication experience in the top 5% of high-impact business and management journals (e.g., Academy of Management Journal, Strategic Management Journal, Administrative Science Quarterly). Performs the manuscript revision workflow: elevates academic voice to a sober, authoritative, and highly critical register; imposes structural framing on every section and sub-section; verifies and completes exhibit citations; sharpens argumentative and statistical exposition; and delivers original-versus-revised edits with editorial notes. Interacts with the reporting and editorial layers.
-
Replication Researcher / Data Reuser. Retrieves the deposited datasets, data dictionary, source register, coding log, and replication code with environment lockfile and README, re-runs the estimation and figure scripts, and extends or contests the study. Interacts with the replication layer.
External recipients (not active system personas): the Financial Regulatory Authority and the EGX; professional associations such as auditors' and governance institutes; top management and boards; institutional and individual investors; and future-research teams addressed by the recommendation set.
System actors: Yahoo Finance (daily equity pricing provider), World Bank Development Indicators (macro series), StockQ (EGX30 annual return series), and EGX/FRA archival sources (issuer filings and investor-relations archives).
Page 9 of 15
5. Core User Flows
Flow 1 — Panel Construction (Empirical Disclosure Researcher).
Define the 2016–2025 EGX non-financial target population → apply the bank/insurance/brokerage/fund/leasing exclusion → apply inclusion criteria (Yahoo ticker history, ≥5 years with ≥20 valid observations, non-missing close/adjusted close/volume, available public archive) → retain the 34-issuer complete universe of continuously archived issuers → record the stage-by-stage exclusion waterfall (Figure 1) → ingest Yahoo Finance daily series and World Bank macro series → assemble 322 firm-year observations (78,037 firm-day observations) → run zero-duplicate-key, zero-blank-cell, and zero-formula-error quality checks → emit the data dictionary and source register.
Flow 2 — Disclosure Index Coding (Empirical Disclosure Researcher).
Retrieve public issuer and investor-relations archives per firm-year → locate each of the four financial items (audited statements, annual report, earnings announcement, ratios/EPS) and each of the three non-financial items (ESG/CSR, corporate governance, board/stakeholder) → record presence/absence with a source URL and retrieval date per cell → compute FDI and NFDI → apply the one-year calendar lag to yield 288 lagged firm-years → apply the prespecified double-coding protocol on the verified subsample and record outstanding κ/ICC as a limitation.
Flow 3 — Estimation and Robustness (Empirical Disclosure Researcher).
Assemble the 284-observation common sample (33 firms) and the 288-observation outcome-specific sample → run ADF unit-root, VIF, Breusch–Pagan LM, Hausman, Wooldridge, and modified Wald diagnostics → estimate the pooled benchmark regressions with industry indicators, year indicators, macro controls, and firm-clustered errors across all four outcomes → estimate the two-way fixed-effects models for all four outcomes with the macro controls absorbed by year effects → report firm-clustered asymptotic p-values, wild-cluster bootstrap p-values, and Benjamini–Hochberg adjusted p-values across the eight-test family → execute the robustness battery (placebo leads, 1%/99% winsorization, hand-verified subsample, leave-one-sector-out, daily extension, Mundlak cross-check) → compile Tables 1–8 and Figures 1–5 → classify the NFDI–return association as suggestive and all estimates as conditional associations.
Flow 4 — Manuscript Editorial Revision (Scholarly Editor / Peer Reviewer).
Ingest the authoritative uploaded manuscript (revised_manuscript_Financial_and_Non-Financial_Disclosure_4.docx) → pass each section for sophisticated academic voice, precise terminology, and removal of redundancy, passive overuse, and colloquialism → scrutinize transitions and section boundaries → impose opening framing and closing synthesis sentences at every section and sub-section → verify every figure and table citation and insert professionally phrased citations where exhibits are implied but unreferenced → strengthen the logical flow of arguments, theoretical models, concepts, and empirical results → verify statistical precision and transparency regarding significance, effect size, and inference caveats → output original text versus revised text for each edit, followed by an "Editorial Notes" block recording structural adjustments such as added transitional sentences and corrected exhibit citations.
Flow 5 — Replication (Replication Researcher / Data Reuser).
Retrieve the deposited derived datasets, data dictionary, source register, coding log, and replication code → restore the environment from the lockfile → re-run the dataset build, estimation, and figure scripts → reproduce Tables 1–8 and Figures 1–5 → confirm that raw issuer reports are not redistributed and that Yahoo Finance terms govern market data → extend the design (exact-date event registry, content-verified indices, MENA cross-country panels).
Page 10 of 15
6. Visuals Colors and Theme
The system's visual layer is a print-first academic exhibit system, not a product interface. Because no palette is specified in the source material, the following restrained defaults are derived from the domain (empirical finance in an elite management journal), the audience (peer reviewers, editors, and replication researchers), and the delivery shape (static tables and figures):
- Base treatment: monochrome grayscale throughout, engineered so that every exhibit remains legible under black-and-white print. Fills use a four-step gray ramp (10%, 30%, 55%, 80%) to encode categories such as FDI, NFDI, and control variables.
- Single accent: one muted, colorblind-safe accent (a desaturated deep teal) reserved exclusively for the primary NFDI–return estimate and its confidence interval in the coefficient plot (Figure 5). No other exhibit uses the accent, so the primary finding reads instantly.
- Negative/positive encoding: direction is carried by position and by line weight, not by red/green semantics; null bands are drawn as a neutral gray reference line.
- Figure family and format: the sample-selection waterfall (Figure 1) as a descending gray step chart; the disclosure-index trajectory (Figure 2) as a two-series line chart on a 0–1 axis with the FDI ceiling at 1.00 marked by a dashed reference rule; the return distribution (Figure 3) as a log-scaled count histogram with a magnified inset over the approximately 96% mass range; group volatility bars (Figure 4) with 95% confidence intervals; and the coefficient plot (Figure 5) on a symmetric-log scale accommodating percentage-point units.
- Typography: a serif face for table and figure body text to match manuscript typesetting, with a sans-serif face for axis and panel labels at one step smaller. Numerals are tabular for column alignment.
- Tables: horizontal rules only, no vertical rules; coefficients and standard errors in aligned columns with clustered standard errors in parentheses, and significance markers reported as nominal, cluster-robust asterisks with the bootstrap and BH-adjusted p-values given in dedicated rows. The pooled benchmark table (Table 5) and the preferred fixed-effects table (Table 6) share this identical vertical grammar so that the benchmark-to-preferred comparison is read structurally rather than narratively.
7. Signature Design Concept
The signature concept is "the source-registered ledger confirmed by the coefficient plot, with the pooled benchmark as the visible foil." Three ideas carry the identity of the system:
- The source register as the visible spine. Every coded cell of FDI and NFDI resolves to a retained source URL, a publication date, and a retrieval date, and every variable resolves to a row-level entry in the data dictionary. The waterfall (Figure 1) makes the numeric reduction from screened population to 284-observation common sample auditable at each stage. This is what distinguishes the artifact from a typical disclosure index: availability is documented, not asserted.
- The pooled benchmark as the deliberate foil. The pooled benchmark regressions (Table 5) are presented in the same reporting grammar as the preferred fixed-effects table (Table 6) precisely so that the reader can see the cross-sectional pattern and then watch it be absorbed by firm and year fixed effects. The benchmark is not filler; it is the visible demonstration of why the preferred specification is necessary.
- The coefficient plot as the single decisive exhibit. Figure 5 consolidates all two-way fixed-effects point estimates and firm-clustered 95% confidence intervals across the four outcomes on one axis, with only the NFDI–return estimate carrying the accent color. The exhibit is engineered so that a reader sees, in one glance, that only non-financial disclosure shifts away from zero against the primary outcome while every other association is bounded near zero.
Together, the concept enforces the study's central epistemic claim: the transparency of the register underwrites the modesty of the inference.
Page 11 of 15
8. Interaction Model & Motion Direction
Because the delivery shape is a static research artifact and print manuscript rather than an interactive application, "interaction" is defined as analytical sequence and exhibit reading order, and "motion" is defined as transition between analytical states. The following direction applies:
- Sequenced inference order. Model selection proceeds as a single ordered pathway: pre-estimation diagnostics → pooled benchmark regressions → preferred two-way fixed-effects → small-cluster and multiple-testing inference → robustness battery. Each step is entered only after the prior step's output is fixed, so that the reader can follow the identification logic without backtracking.
- Benchmark-to-preferred transition. The pooled benchmark regressions serve as the first estimation state; the argument then moves to the fixed-effects state, where the benchmark's cross-sectional coefficients are visibly absorbed. The transition is framed as a deliberate narrowing of identification, not a duplication of results.
- Evidence-to-claim ordering in the text. Each exhibit is introduced before the claim it supports, and every claim resolves to a named exhibit; narrative transitions carry the reader from unconditional description to conditional estimation to bounded interpretation.
- Regression-table reading rhythm. Each table follows a fixed vertical grammar: coefficient row, standard error in parentheses, asymptotic cluster-robust t and p, wild-cluster bootstrap p, Benjamini–Hochberg adjusted p, then fixed-effect and observation counts. The repeated rhythm is the table-level analogue of motion, and its consistency across the pooled and fixed-effects tables is what makes the comparison legible.
- Descending-to-decisive exhibit order. The document's visual argument moves from breadth to focus: waterfall → index trajectory → distribution → group descriptives → coefficient plot. The sequence contracts from population to single estimate.
- No motion in the artifact. Figures are static, animate nothing, and rely on position, weight, and the single accent to direct attention. No transitions, hover states, or interactive controls are introduced, because the artifact must survive print and archival deposit unchanged.
- Interpretive damping as a deliberate direction. Every transition from estimate to meaning narrows rather than amplifies: nominal significance yields to bootstrap significance, which yields to BH-adjusted significance, which yields to the "suggestive rather than confirmatory" characterization. The motion of the argument is toward restraint.
Page 12 of 15
9. Non-Functional Requirements
- Reproducibility. The complete chain from raw sources to reported estimates must be re-executable from the deposited dataset build, estimation, and figure scripts together with an environment lockfile and README. No estimate may depend on undocumented manual steps.
- Data provenance. Every coded disclosure cell must be traceable to a retained source URL and retrieval date; every variable must be traceable to a data-dictionary and source-register entry.
- Statistical validity. Stationarity must be established before inference; multicollinearity must clear conventional VIF thresholds; panel specification must be tested and defended; small-cluster inference must be supplemented with bootstrap p-values; and the pre-specified family of eight tests must be reported with Benjamini–Hochberg adjustment.
- Transparency of uncertainty. Nominal, cluster-robust, wild-bootstrap, and BH-adjusted p-values must be labeled distinctly wherever they appear, and the primary result must be characterized as suggestive rather than confirmatory.
- Interpretive discipline. All estimates must be described as conditional associations, with no causal language, and the competing explanations for the null FDI result must be left undiscriminated.
- Internal validity of the panel. Zero duplicate firm-date keys, zero blank cells, and zero formula errors must be confirmed; the sample-rule choice (common-sample versus outcome-specific) must be disclosed per equation.
- Comparability across specifications. The pooled benchmark regressions and the preferred fixed-effects models must be reported with parallel table structure, variable ordering, and inference rows, so that the benchmark-to-preferred comparison is structural rather than interpretive.
- Auditability of exclusion. The inclusion/exclusion waterfall must report stage-by-stage counts, and the external-validity boundary must be stated.
- Measurement honesty. The archive-availability limitation, the coarse index support, the proxy share-count denominator, the calendar-versus-fiscal-year alignment risk, and the outstanding κ/ICC statistics must be reported as residual risks rather than silently mitigated.
- Ethics and data rights. The study must declare use of public secondary data with no human participants; raw issuer reports must not be redistributed where copyright restricts it; Yahoo Finance data must remain subject to provider terms.
- Journal compliance. Every reference must carry a functioning DOI, all bracketed publication details must be resolved, and the 2024–2026 literature base must be expanded to approximately twenty studies.
- Editorial consistency. The manuscript must maintain a uniform authoritative academic register, complete exhibit citation coverage, and a consistent original-versus-revised plus Editorial Notes output format across all edits. Every section and sub-section must open with a framing sentence and close with a synthesizing bridge sentence, and the editorial revision must preserve the manuscript's substantive content and meaning.
- Archival stability. All exhibits must remain legible and unchanged when reproduced in grayscale print and in archival deposit.
Page 13 of 15
10. Tech Stack
The accepted delivery shape is a self-contained empirical research artifact with a deposited replication package. Only the layers that shape actually requires are specified; no front end, authentication, administration, or service layer is included.
- Analysis language and runtime: Python 3.12.
- Core libraries: pandas, statsmodels, linearmodels.
- Market data source: Yahoo Finance daily equity pricing series (close, adjusted close, volume).
- Macroeconomic data source: World Bank Development Indicators — inflation (FP.CPI.TOTL.ZG), official EGP/USD rate (PA.NUS.FCRF), GDP growth (NY.GDP.MKTP.KD.ZG).
- Market benchmark source: EGX30 annual return series (StockQ, 2026).
- Disclosure source: public issuer filings and investor-relations archives, with retained source URLs, publication dates, and retrieval dates.
- Estimation output and replication repository: deposited derived annual and daily datasets, data dictionary, source register, coding log, figure and estimation scripts, environment lockfile, README, and the accompanying workbook. The estimation scripts cover both the pooled benchmark regressions and the two-way fixed-effects models.
- Manuscript production: static typeset tables and figures suitable for grayscale journal reproduction, and an editorial revision workflow applied to the authoritative uploaded manuscript (revised_manuscript_Financial_and_Non-Financial_Disclosure_4.docx) that delivers each edit as original text versus revised text followed by an "Editorial Notes" block.
Page 14 of 15
11. Assumptions and Constraints
- Associational design. The study estimates conditional associations only; no instrument, regulatory discontinuity, or exact announcement dates are available, so causal interpretation is excluded.
- Archive availability, not content. FDI and NFDI record whether items were located in cited public archives; they do not measure disclosure quality, timeliness, or content. A zero denotes "not located," not verified non-disclosure.
- Benchmark status of pooled estimation. The pooled benchmark regressions are a comparison baseline, not the preferred specification; the pooled coefficients are interpreted only to the extent needed to motivate the two-way fixed-effects framework, with no inference claimed from the benchmark alone.
- Joint identification constraint. Because annual macro variables are common within a year, the pooled design (which identifies them) and the fixed-effects design (which absorbs them) are reported as alternatives and cannot be simultaneously identified.
- Sample composition. The 34 firms constitute the complete universe of eligible non-financial EGX issuers with surviving, continuous public archive data, not a random cross-section; results generalize only to continuously archived issuers.
- Saturation asymmetry. FDI reaches 1.00 for all sampled issuers by 2020, confining within-firm FDI variation largely to the pre-2020 period and muting the identifying content of financial-archive items.
- Coarse index support. FDI takes only the values 0.25 and 1.00; NFDI only 0.33, 0.67, and 1.00. Policy-like jumps around 2020–2022 behave partly as common-shock adoption indicators.
- Small clusters. The preferred annual models rely on 33 firm clusters; the document-verified subsample contains only four firms, from which no inference beyond directional attenuation-bias interpretation is drawn.
- Outstanding reliability statistics. Cohen's κ and ICC for the double-coding protocol remain outstanding and are reported as a limitation.
- Proxy denominators. Current share-count denominators may misstate historical turnover and capitalization; turnover is therefore demoted to a secondary outcome and log volume is primary for trading activity.
- Timing alignment. Calendar-year rather than fiscal-year alignment is used for some issuers (notably EAST), which risks timing misalignment and attenuation of lagged associations.
- Sample-rule dependence. The common-sample rule yields 284 observations across 33 firms; the outcome-specific rule retains 288 for return and volatility. Rule choice affects comparability across equations and is disclosed per table, in both the pooled benchmark and fixed-effects tables.
- Daily extension carries no independent identifying information because annual indices are carried forward to the firm-day level.
- Data rights. Raw issuer reports are not redistributed where copyright restricts it; Yahoo Finance data remain subject to provider terms.
- Reference compliance. Every reference must carry a functioning DOI, and publication details currently bracketed as to-be-verified must be resolved before submission.
- No event-study estimates are reported because exact announcement dates are not available; an exact-date event registry is identified as a priority for future research.
- Editorial scope. The editorial workflow revises prose, structure, exhibit citation coverage, and argumentative exposition only; it must not alter the manuscript's substantive content, empirical results, numbers, or claims, and it operates on the authoritative uploaded manuscript as its content source.
Page 15 of 15
12. Glossary
- FDI (Financial Disclosure Index): A four-item archive-availability index = (audited statements + annual report + earnings announcement + ratios/EPS)/4, lagged one year.
- NFDI (Non-Financial Disclosure Index): A three-item archive-availability index = (ESG/CSR + corporate governance + board/stakeholder)/3, lagged one year.
- Archive availability: Whether a specific disclosure item was located in cited public issuer or investor-relations archives; a measure of availability, not of quality, timeliness, or content.
- Pooled benchmark regressions: The pooled OLS specifications estimated with industry indicators, year indicators, macroeconomic controls, and firm-clustered standard errors across all four outcomes, reported as the comparison baseline to the preferred two-way fixed-effects models.
- Pooled OLS: Ordinary least squares estimation on stacked firm-year observations without firm-specific effects, used here only as a benchmark.
- EGX: The Egyptian Exchange.
- FRA: The Financial Regulatory Authority (Egypt), which tightened periodic-reporting enforcement after 2016 (Decrees No. 107 and 108 of 2021).
- ORAS: A sampled issuer with four firm-years (2022–2025) recording zero annual trading volume, for which log volume is undefined.
- Annual adjusted return: Percentage change between the first and last valid adjusted closing prices per firm-year.
- Annualized daily volatility: Within-year standard deviation of daily log returns × √252.
- Turnover ratio: Annual volume divided by the public share-count denominator.
- Two-way fixed effects: A panel specification including both firm and year fixed effects, absorbing time-invariant firm heterogeneity and common macroeconomic shocks.
- Firm-clustered standard errors: Standard errors clustered at the firm level, robust to within-firm correlation and heteroskedasticity.
- Wild-cluster bootstrap: A Rademacher bootstrap procedure (999 replications) generating p-values that remain reliable when the number of clusters is small.
- Benjamini–Hochberg (BH) adjustment: A multiple-testing procedure controlling the false-discovery rate across the pre-specified family of eight hypothesis tests.
- Mundlak / correlated-random-effects specification: A random-effects model including cluster means of the regressors, used as a cross-check on the fixed-effects choice.
- Pre-specified family of eight tests: Two disclosure indices × four market outcomes, with annual adjusted return designated the primary outcome.
- ADF / IPS panel unit-root screen: Firm-level augmented Dickey–Fuller tests on daily log returns constituting an Im–Pesaran–Shin-style stationarity screen.
- VIF: Variance inflation factor, used to screen multicollinearity.
- Breusch–Pagan LM test: A Lagrange Multiplier test of pooled OLS against random effects.
- Hausman test: A specification test discriminating fixed from random effects.
- Wooldridge test: A test for first-order serial correlation in panel data.
- Modified Wald test: A test for groupwise heteroskedasticity in fixed-effects panels.
- κ / ICC: Cohen's kappa and intraclass correlation, reliability statistics for the prespecified double-coding protocol; currently outstanding.
- CAR: Cumulative abnormal return, a future-research event-study measure requiring an exact-date disclosure registry.
- ESG: Environmental, social, and governance disclosure content.
- ISSB / IFRS S1–S2: International Sustainability Standards Board standards referenced in the recommendation set.
- CSRD: The EU Corporate Sustainability Reporting Directive, referenced in the literature base.
- EGX30: The Egyptian Exchange headline index; its annual return series enters the pooled benchmark regressions as a macro control.
- MENA: Middle East and North Africa, the regional extension frame for future research.
- Conditional association: A statistical relationship estimated holding specified controls and fixed effects constant, carrying no causal claim.
- Suggestive finding: A nominally significant association that does not survive multiple-testing adjustment and is therefore not characterized as confirmatory.
- Editorial Notes: A brief block placed below each revised passage recording the specific structural adjustments made to that passage, such as added transitional sentences or corrected exhibit citations.
- Original-versus-revised presentation: The required editorial output format in which each edit is shown as the original text alongside the revised text, so that every change is auditable.
No comments yet. Be the first!