Section 01

Guiding principles

Six rules apply to every piece of work covered by this document. They are not aspirational; they are constraints. Work that violates one of them is not eligible for promotion, regardless of the numbers it produces.

  1. Data integrity before component quality. A component trained on data that would not have been available at prediction time is not a component. It is a curve fit. Point-in-time correctness is the first gate, not the last.
  2. Baselines before models. No predictive component is evaluated against a null. Historical mean, persistence, and zero-return baselines are computed for every target before any non-trivial model is fit. A model that does not beat its baselines out-of-sample has no claim on being a model.
  3. Chronology, not randomness. Time series are never split with train_test_split or equivalent random shuffling. Splits are chronological, with purge and embargo applied where horizons overlap.
  4. Each component is validated against its own question. A state engine, a forecast model, a retrieval system and a decision layer are validated against different evidence. No component's validation is evidence for another's, and no track gates the others.
  5. Failures are published. A validation report is only meaningful if it discloses what did not work. Failed windows, unstable coefficients, non-converged fits, and rejected approaches are recorded alongside the successful results.
  6. Reproducibility is a property of the pipeline, not the notebook. Every experiment must be reproducible from an identifier plus the recorded dataset, feature, code, and configuration versions. Anything less is not research; it is exploration.

Section 02

Validation by component class

Tickflow Capital validates every component against the question it was built to answer. There are two independent tracks - strategy validation and analytical validation - and within analytical validation there are several component classes, each with its own evidence type. No track gates another. No class's validation is evidence for another's.

Track A

Strategy validation

Does this trading rule have a positive expected edge on unseen data, after realistic execution assumptions?

Primary evidence
  • Walk-forward windows on unseen validation data
  • Profit factor per window with pass/fail criterion
  • Forward-period returns on an untouched test window
  • Maximum drawdown, Sharpe, Sortino, Calmar
  • Trade-level statistics including trade count and win rate
  • Monte Carlo confidence intervals on the return distribution
  • Multiple-testing adjustment across the window set
Consumed by

The trading account. Deployment of capital. No analytical component is authorised to claim Track A validation as evidence of its own correctness.

Track B

Analytical validation

Does this component answer the analytical question it was built to answer, on unseen data, with the appropriate stability and calibration for its class?

Primary evidence

Depends on component class - see the table below. ML models are validated for out-of-sample performance and calibration; statistical engines for estimate stability; retrieval systems for relevance; event studies for sample adequacy; decision layers for rule-outcome alignment; the interface for routing accuracy and evidence traceability.

Consumed by

The analytical platform, downstream research, and the components that depend on this one. Never consumed by the trading account.

Component classes within Track B

Each class is validated against its own evidence type. Applying the wrong evidence to the wrong class is itself a validation failure.

Component class Examples in the platform Primary validation evidence
Infrastructure Point-in-time data layer, feature engine, data quality engine Determinism under replay, vintage integrity, schema stability
Statistical engines Macro state engine, change engine, relationship engine, statistical laboratory Coefficient and estimate stability under bootstrap resampling; correlation of derived quantities with realised outcomes
ML models Macro regime, XAU direction, XAU distribution, XAG direction, XAG distribution Out-of-sample likelihood or R², calibration, cluster stability, persistence over time
Retrieval systems Historical analogue engine Similarity relevance; forward-outcome distributional stability across retrieved episodes
Event studies Event reaction engine Sample adequacy per event type; conditional distribution stability; coverage of event categories
Simulation engines Scenario engine Propagation correctness under shock; calibration of resulting distributions; replay verification
Decision layers Thesis engine, position state engine, risk engine Rule-outcome alignment; invalidation discipline; drawdown response consistency
Interface Question engine Routing accuracy; evidence traceability; no hallucinated numbers in synthesised answers
A third track exists in principle, but not yet in practice. If a strategy is ever modified to consume an analytical component's output as an input, that integration requires its own validation (Track C), governed by Track A's evidence standard. Until then, no Track B result can be cited as evidence for a Track A claim, or vice versa. The Macro Intelligence Platform is architecturally independent from the trading engine, and this document does not conflate them.

Section 03

Point-in-time data integrity

The single most common source of inflated backtest performance in macro and financial ML is the use of revised data. This section describes how Tickflow Capital prevents that class of error.

The requirement

For every row of training data, the pipeline must be able to demonstrate that every feature value used was published on or before the prediction timestamp. This is not a property of the code; it is a property of the storage layer. Data is stored with an available_at field, and every query filters on available_at <= as_of.

Vintage handling

Macroeconomic series are revised. GDP, employment, and inflation series are subject to multiple revisions after first publication. Using today's value of a series to backtest a decision made two years ago is look-ahead, even if no future-dated observations are used. Every series in the pipeline is graded:

Vintage grade Meaning Eligibility
Exact Historical first-release values are preserved and retrievable for each publication date. Eligible for model features and backtesting.
Approximate Some vintage history exists but is incomplete, or has been reconstructed from revision logs. Eligible with a documented accuracy bound. Must be disclosed in the validation report.
Unavailable Only the latest revised value is retrievable. Historical vintages do not exist in any accessible source. Not eligible for model features used in backtesting. May be used for the live dashboard, with an explicit point-in-time caveat.
Blocking issue. A model feature set is not considered complete until every input has been assigned a vintage grade, and every "unavailable" series has either been dropped or replaced. This is currently the gating item for the regime model's first fit.

Data quality gates

Before a dataset is eligible for training, the following checks must pass. Any failure halts the pipeline and produces a machine-readable reason.

Schema conformity. Every series matches its expected schema: column names, types, frequency, and units.
Timestamp ordering. Observation, release, and availability timestamps are monotonic and consistent with the source's publication schedule.
Duplicate detection. No two rows claim the same series, observation date, and vintage.
Gap analysis. Missing observations within a series are enumerated and classified (weekend, holiday, delayed publication, data error).
Stale series detection. Series whose last observation is older than their expected refresh interval are flagged.
Future-date rejection. No observation may carry an available_at later than the dataset's stated availability cutoff.
Unit-change detection. Series that change units or frequency mid-history are flagged and require explicit handling.
Provider disagreement. Where the same conceptual series is available from multiple providers, values are compared and material differences flagged.

Section 04

Split design and leakage prevention

Time-series data cannot be randomly split. Adjacent observations are correlated, targets overlap, and a randomly chosen test set contains rows whose information was already visible to the model during training.

Chronological splits

Every dataset is split chronologically. The default structure is train–validation–test, with each split covering a contiguous block of time and no overlap between them. Expanding-window and rolling-window walk-forward are supported and preferred for final evaluation.

Purge and embargo

When the target is a forward return over horizon h, the last h rows of the training set have targets that extend into the validation set. Those rows must be purged. Immediately after each split boundary, an embargo period is applied so that the first rows of the validation set do not share information with the last rows of the training set.

Training Purge last h rows Embargo gap Validation Purge last h rows Embargo gap Test

The test set is touched once

The final test set is not consulted during model development, feature selection, hyperparameter tuning, or any other iterative process. Consulting it more than once per artifact version converts it into a validation set, and invalidates the out-of-sample claim.

Leakage tests

Every pipeline runs an automated leakage test suite. A test failure is a hard block, not a warning.

Feature timestamp ≤ prediction timestamp for every row.
Feature available_at ≤ prediction timestamp for every row, including vintage dates.
Target horizon purge applied at every split boundary.
Scaler and normalizer fit on training only, then applied to validation and test.
Label encoders built from training categories only, unknown categories at inference treated as missing.
No target-derived feature enters the model. Rolling statistics computed on features must not include the target.
Shuffle test. Shuffle the target and re-fit. If the model improves on the shuffled target above chance, a leak exists.

Section 05

Baselines before models

Every predictive target is evaluated against a set of trivial baselines before any non-trivial model is fit. A model that does not beat its baselines out-of-sample has no claim on being a model.

Standard baselines

Baseline Prediction Purpose
Historical mean Training-set mean of the target Establishes whether the model adds information beyond the unconditional expectation.
Historical median Training-set median of the target Same as above, more robust to outliers.
Persistence / zero Zero forward return The default for a strategy that does not trade. A model must beat not trading.
Most recent value Last observed target value Tests whether the model captures any momentum or mean reversion.
The baseline comparison is not optional. Validation reports list baseline metrics alongside model metrics, side by side. A model that reports only its own metrics is not a completed experiment.

Section 06

Validation gates

Every component must clear a set of gates before it can be promoted between lifecycle stages. Gates are hard thresholds. Where a threshold is not yet defined, the gate is listed as pending and no promotion is possible until it is defined. The full gate set is applicable to components that produce predictions; infrastructure and retrieval components have a reduced gate set defined by their class.

Data gates

Gate Threshold Evidence required
Point-in-time correctness 100% of rows pass available_at ≤ as_of Leakage test suite run log
Vintage coverage 100% of features have a vintage grade assigned Feature registry with vintage column populated
Schema conformity Zero schema errors at ingest Data validation report
Sample size Threshold pending per component class Row count after all filters

Statistical gates

Gate Threshold Evidence required
Beats baselines out-of-sample Must exceed all four standard baselines on primary metric Side-by-side metrics table
Out-of-sample vs. in-sample gap OOS performance within a documented fraction of IS performance Overfitting diagnostic report
Multiple-testing adjustment p-values adjusted for the number of configurations tested Configuration count and adjustment method
Calibration Threshold pending per output type Reliability diagram, Brier score or equivalent

Stability gates

Gate Threshold Evidence required
Cluster or coefficient stability Stable under bootstrap resampling Bootstrap diagnostic, at least 200 resamples
Persistence over time Every defined state or cluster persists for a documented minimum Assignment timeline
Sub-period stability Performance not concentrated in a single historical regime Segmented performance report
Feature stability Coefficient signs and importances do not flip under resampling Feature diagnostics

Coherence gates

Gate Threshold Evidence required
Economic interpretability Every cluster centroid, coefficient, or state transition is describable in plain English Written interpretation accepted by review
Reference-set agreement Agreement with hand-labelled historical periods at a rate above chance Confusion matrix against the reference set
Dimensional consistency Every composite has a single semantic axis with all components aligned in sign Feature registry with sign convention documented
Where a threshold is marked "pending", no promotion occurs. A gate that has not been defined is not a gate. It is a placeholder. Components whose validation reports contain pending gates remain in the Research stage.

Section 07

Reproducibility and experiment tracking

Every experiment must be reproducible from an identifier plus the recorded inputs. If the same experiment is run twice on the same dataset, it must produce the same outputs. If it does not, the discrepancy is a bug and must be resolved before the experiment is treated as evidence.

What every experiment records

Experiment ID - stable identifier, unique across the registry.
Dataset ID - exact snapshot used, with row count and date range.
Feature version - identifier of the feature set used, with the versioned registry entry it resolves to.
Target version - identifier of the target construction, including horizon.
Model type and hyperparameters - full configuration, not just the class name.
Random seed - for every source of randomness, including initialisation and resampling.
Train, validation, test boundaries - including purge and embargo settings.
Code commit - git SHA of the repository at run time.
Dependency lockfile - exact versions of every library in the environment, not just the top-level names.
Environment - Python version, platform, and any other runtime-affecting detail.
Metrics - every computed metric, per split, in a machine-readable form.
Artifact hash - sha256 of the trained model, scaler, and feature schema, so the exact artifact is verifiable.
Artifact path - location of the trained model, scaler, and feature schema in the repository or model store.
Last retrain timestamp - when the current artifact was last produced, so staleness is visible.
Monitoring status - which drift signals are being tracked against this artifact, and their current state.
Status - current lifecycle stage and promotion history.

The reproducibility contract

An experiment is considered reproducible if, given the experiment ID and the code commit, a full re-run produces:

  1. The same dataset snapshot, with the same row count and date range.
  2. The same feature matrix, element for element.
  3. The same model configuration, including fitted parameters, to within the documented numerical tolerance.
  4. The same out-of-sample metrics, to within the documented numerical tolerance.
  5. The same artifact hash, or a documented reason why the hash differs (e.g. platform-level floating-point divergence).
  6. The same validation report, modulo timestamps.
Reproducibility is a gate, not an aspiration. A component whose experiments cannot be reproduced does not pass the reproducibility gate and remains in Research regardless of its metrics.

Section 08

Lifecycle and decision authority

Every component moves through a defined lifecycle. Promotion between stages is a deliberate act, requires the evidence appropriate to the target stage, and is recorded permanently in the registry. In addition to lifecycle stage, every component carries a decision authority - the maximum scope of action its output is permitted to inform. Lifecycle and authority are separate axes; a component can be in Production and still be scoped to context-only authority.

Lifecycle stages

Design Research Candidate Validated Shadow Production Monitored Degraded Retired
Stage Requirements to enter Consumed by
Design Specification complete. The question the component answers, its inputs, its outputs and its dependencies are documented. Implementation has not begun. Internal planning only.
Research Implementation in progress. Feature engineering, hypothesis testing, offline evaluation. No external output. Internal only.
Candidate Dataset, feature version, code commit, lockfile and configuration recorded. Artifacts committed to the repository. Internal only.
Validated All gates applicable to the component class passed. Validation report published internally. Human sign-off recorded. Internal research. Eligible for shadow deployment.
Shadow Minimum period of live-data operation with predictions compared to outcomes. No capital exposure, no live decision influence. Internal research. Publicly listed as shadow.
Production Shadow metrics pass defined thresholds. No regressions. Explicit sign-off. Decision authority assigned. Live use, scoped by decision authority.
Monitored Continuous production operation with all drift signals tracked, alerting active, and periodic re-evaluation scheduled. Live use, unchanged. Distinguished from Production by the completeness of monitoring.
Degraded Live performance departs from shadow expectations beyond documented tolerance, or a monitoring signal breaches its threshold. Output retained for analysis. Not used for decisions.
Retired Superseded, decommissioned, or found to be invalid. Historical reference only.
No automatic promotion, ever. The registry accepts promotion commands only from an authenticated human action. There is no scheduled job, CI rule, or automated process that moves an artifact between stages.

Decision authority

Every component carries a decision authority, independent of its lifecycle stage. A component in Production with Context authority cannot inform a position; a component in Shadow with Decision-support authority is not yet permitted to act on it. Authority is assigned explicitly and recorded in the registry.

Authority Meaning Consumed by
Context Informs human reading. Makes no recommendation. Humans reading the dashboard.
Analytical Produces outputs consumed by other analytical components. Downstream components.
Forecast Produces forward-looking estimates with calibrated uncertainty. Downstream components and humans.
Decision-support Produces recommendations consumed by humans. Humans deciding whether to act.
Execution Authorised to affect positions or capital. The trading account.
Lifecycle and authority are separate axes. A component's stage says how mature its work is. Its authority says what its output is permitted to inform. The two are recorded independently in the registry so that a production component with context-only authority is not mistaken for one authorised to drive decisions.

Section 09

Shadow deployment

A component that has cleared its validation gates has proven itself on historical data. It has not proven itself on live data. Shadow deployment is the step between the two.

Evidence types — what "simulated" means

Five kinds of evidence appear at different stages of a component's life, and they are not interchangeable. Where the registry or the applied record describes a figure as simulated, it was produced by one of the first three categories below — not by live trading. Only the final category constitutes a live track record.

Evidence type What it involves What it proves What it does not prove
In-sample backtest Strategy is fit and evaluated on the same historical data. That parameters exist which fit the given history. Nothing about future performance. Reported for completeness only.
Walk-forward validation Strategy is repeatedly fit on a training window and evaluated on the immediately following unseen window, rolled forward. That the edge, if it exists, is not a product of parameter selection on the tested period. That the edge will persist. Each window is an independent trial, not a live record.
Simulated holdout A single period is held back from all training and parameter selection, and evaluated once. The resulting figures are backtest output. That the strategy generalises to a period the model never saw, under stated execution assumptions. That the strategy has been deployed, that capital was at risk, or that real execution costs match the assumptions.
Shadow deployment The component runs against live data on schedule, producing outputs that are stored but not acted on. That the component operates correctly on live data, with outputs compared against real outcomes. That the component has influenced positions or capital.
Live track record The component or strategy is deployed and operating on real capital, with real execution. That the strategy performs under live conditions, with all associated frictions. That it will continue to perform.

Where the applied record labels a period simulated holdout, that label states which category of evidence the figures belong to. It is not a claim of live trading. The registry uses this classification when assigning lifecycle stages: a component that has passed walk-forward and holdout evaluation but has not been deployed is marked Validated, not Live.

What shadow means

A shadow component runs against live data on a fixed schedule. It produces outputs on the same cadence as the production system would. Those outputs are stored, timestamped, and preserved alongside the input features that produced them. They are not used to make decisions.

What is measured

Output quality

  • Whether outputs match the eventual observed outcomes
  • Whether calibration holds out-of-sample
  • Whether uncertainty intervals contain the realised values at the expected rate
  • Whether the distribution of outputs is consistent with the training-time distribution

Operational quality

  • Whether the component runs on schedule without errors
  • Whether its inputs are fresh and complete
  • Whether its outputs are computable within the required latency
  • Whether any data-quality warnings were raised during the shadow period

Duration and exit

Shadow duration is set by the component's output cadence and, where applicable, the target horizon. A monthly regime classifier requires a longer shadow period than a daily directional model because fewer outcomes are observed per unit time. Exit criteria are defined before shadow deployment begins; they are not adjusted mid-period to accommodate the observed results.

Exit criteria are set in advance. Adjusting shadow exit criteria after seeing shadow performance is a form of overfitting. If the criteria need revision, the shadow period restarts.

Section 10

Monitoring and drift

A component in production is not assumed to remain valid indefinitely. The following are monitored continuously, and departures trigger defined responses.

Signal What is measured Response on breach
Data freshness Age of each input series relative to its expected refresh interval Flag stale inputs. Suppress component output if critical inputs are stale beyond tolerance.
Input distribution drift Population stability index or equivalent on each feature Record drift event. Escalate to research if sustained.
Output distribution drift Distribution of component outputs over a rolling window Compare against training-time distribution. Flag if materially different.
Calibration drift Reliability of probabilistic outputs against realised outcomes Recompute calibration on recent window. Downgrade if calibration fails.
Performance drift Outcome metric against expected value from validation Compare against shadow-period baseline. Trigger review if outside tolerance.
Operational failures Error rate, latency, retry count Standard operational response. Component output suppressed if failures exceed tolerance.
Degradation is a defined state, not a failure. A component whose live performance departs from its validation expectations enters the Degraded stage. Degraded components continue to produce outputs for analysis, but those outputs are not used for decisions until the component is re-validated or retired.

Section 11

What is published, what is withheld

This page, and the system registry it governs, are intended to be substantive. They are not exhaustive. The following boundary applies to every public document.

Published

  • The methodology itself: gates, splits, baselines, tracking practice
  • The existence of a regime filter, in general terms
  • The count and outcomes of walk-forward windows
  • Forward-period summary statistics
  • The identity of the canonical composites used as features
  • The lifecycle stage and decision authority of each component
  • Failures, corrections, and rejected approaches

Withheld

  • Numeric values of any threshold, cutoff, or parameter
  • Stop distances, entry bands, session blocks
  • Position-sizing inputs and risk-budget constants
  • Feature weights, scoring coefficients, sign conventions
  • Exact dataset composition
  • Internal artifact identifiers beyond registry entries
  • Specific model hyperparameters
The boundary is deliberate. Transparency about process and outcomes is compatible with discipline about parameters. Both are required to run a research firm whose work is both credible and defensible.

Section 12

Rejected approaches

A validation methodology is only as strong as the practices it rejects. The following approaches have been considered and are not permitted in any pipeline governed by this document.

Random train/test splits on time-series data. Adjacent observations are correlated; a random split leaks information from validation into training.
Fitting scalers on the full dataset. Scalers and normalizers must be fit on training data only and then applied unchanged to validation and test.
Reporting in-sample performance as evidence. A component's fit to data it was trained on is not a validation result.
Ignoring multiple-testing effects. Testing 50 configurations and reporting the best one without adjustment is a form of p-hacking.
Using revised macro data without vintage awareness. As described in Section 03, this invalidates any point-in-time claim.
Auto-promotion from validation to production. Every promotion is a deliberate, recorded human action.
Publishing validation claims before they are verified. A registry entry is a claim. Claims that have not been verified are not published.
Publishing operational history that did not occur. Deployment dates, evaluation dates and lifecycle transitions must reflect what actually happened. A status page that shows events on a timeline is not permitted to invent those events.
Adjusting thresholds after seeing results. Thresholds are set before validation begins. Changing them mid-validation invalidates the validation.
Treating one track's validation as evidence for another. Track A and Track B are independent. Neither validates the other.
Applying the wrong evidence type to a component class. A retrieval system is not validated by out-of-sample R². A statistical engine is not validated by walk-forward profit factor. Each class has its own evidence standard.
Silent corrections. When an entry is wrong, the correction is recorded, not silently overwritten.

Section 13

Governance

This document is a working standard. It is reviewed quarterly, and any change to it is recorded in the system registry's changelog.

Ownership

This methodology is owned by the firm, not by any individual. The founder is bound by it in the same way as any future team member. No component is promoted without the evidence required by the sections above, regardless of who is doing the work.

Challenging a result

Any validation claim may be challenged. A challenge is a request for the raw data, the configuration, the code commit, the lockfile, and the metrics that support the claim. If those cannot be produced, the claim is withdrawn until they can. This applies to claims made in public documents, in private reports, and in internal discussions.

Response to failure

When a validation claim turns out to be wrong, the response is:

  1. Record the correction in the changelog, dated and attributed.
  2. Determine whether the error was in the evidence, the interpretation, or the process.
  3. If the process is the source of the error, amend this document.
  4. Re-validate any component that relied on the withdrawn claim.
A methodology is only credible if it is applied when the results are inconvenient. Rules that are dropped when they produce bad outcomes are not rules. This document exists so that the temptation to make an exception is visible when it arises.

Questions about this methodology

If you are evaluating Tickflow Capital as a research partner, or you need to understand how a specific result was produced, the contact form goes directly to the founder.

Contact form View system registry Applied results