Section 01
Guiding principles
Six rules apply to every piece of work covered by this document. They
are not aspirational; they are constraints. Work that violates one of
them is not eligible for promotion, regardless of the numbers it
produces.
-
Data integrity before component quality.
A component trained on data that would not have been available at
prediction time is not a component. It is a curve fit.
Point-in-time correctness is the first gate, not the last.
-
Baselines before models.
No predictive component is evaluated against a null. Historical
mean, persistence, and zero-return baselines are computed for every
target before any non-trivial model is fit. A model that does not
beat its baselines out-of-sample has no claim on being a model.
-
Chronology, not randomness.
Time series are never split with train_test_split
or equivalent random shuffling. Splits are chronological, with
purge and embargo applied where horizons overlap.
-
Each component is validated against its own question.
A state engine, a forecast model, a retrieval system and a decision
layer are validated against different evidence. No component's
validation is evidence for another's, and no track gates the others.
-
Failures are published.
A validation report is only meaningful if it discloses what did
not work. Failed windows, unstable coefficients, non-converged
fits, and rejected approaches are recorded alongside the
successful results.
-
Reproducibility is a property of the pipeline, not the
notebook.
Every experiment must be reproducible from an identifier plus the
recorded dataset, feature, code, and configuration versions.
Anything less is not research; it is exploration.
Section 02
Validation by component class
Tickflow Capital validates every component against the question it
was built to answer. There are two independent tracks - strategy
validation and analytical validation - and within analytical
validation there are several component classes, each with its own
evidence type. No track gates another. No class's validation is
evidence for another's.
Track A
Strategy validation
Does this trading rule have a positive expected edge on unseen
data, after realistic execution assumptions?
Primary evidence
- Walk-forward windows on unseen validation data
- Profit factor per window with pass/fail criterion
- Forward-period returns on an untouched test window
- Maximum drawdown, Sharpe, Sortino, Calmar
- Trade-level statistics including trade count and win rate
- Monte Carlo confidence intervals on the return distribution
- Multiple-testing adjustment across the window set
Consumed by
The trading account. Deployment of capital. No analytical
component is authorised to claim Track A validation as evidence
of its own correctness.
Track B
Analytical validation
Does this component answer the analytical question it was built
to answer, on unseen data, with the appropriate stability and
calibration for its class?
Primary evidence
Depends on component class - see the table below. ML models are
validated for out-of-sample performance and calibration;
statistical engines for estimate stability; retrieval systems
for relevance; event studies for sample adequacy; decision layers
for rule-outcome alignment; the interface for routing accuracy
and evidence traceability.
Consumed by
The analytical platform, downstream research, and the components
that depend on this one. Never consumed by the trading account.
Component classes within Track B
Each class is validated against its own evidence type. Applying the
wrong evidence to the wrong class is itself a validation failure.
| Component class |
Examples in the platform |
Primary validation evidence |
| Infrastructure |
Point-in-time data layer, feature engine, data quality engine |
Determinism under replay, vintage integrity, schema stability |
| Statistical engines |
Macro state engine, change engine, relationship engine, statistical laboratory |
Coefficient and estimate stability under bootstrap resampling; correlation of derived quantities with realised outcomes |
| ML models |
Macro regime, XAU direction, XAU distribution, XAG direction, XAG distribution |
Out-of-sample likelihood or R², calibration, cluster stability, persistence over time |
| Retrieval systems |
Historical analogue engine |
Similarity relevance; forward-outcome distributional stability across retrieved episodes |
| Event studies |
Event reaction engine |
Sample adequacy per event type; conditional distribution stability; coverage of event categories |
| Simulation engines |
Scenario engine |
Propagation correctness under shock; calibration of resulting distributions; replay verification |
| Decision layers |
Thesis engine, position state engine, risk engine |
Rule-outcome alignment; invalidation discipline; drawdown response consistency |
| Interface |
Question engine |
Routing accuracy; evidence traceability; no hallucinated numbers in synthesised answers |
A third track exists in principle, but not yet in practice.
If a strategy is ever modified to consume an analytical component's
output as an input, that integration requires its own validation
(Track C), governed by Track A's evidence standard. Until then, no
Track B result can be cited as evidence for a Track A claim, or vice
versa. The Macro Intelligence Platform is architecturally independent
from the trading engine, and this document does not conflate them.
Section 03
Point-in-time data integrity
The single most common source of inflated backtest performance in
macro and financial ML is the use of revised data. This section
describes how Tickflow Capital prevents that class of error.
The requirement
For every row of training data, the pipeline must be able to demonstrate
that every feature value used was published on or before the
prediction timestamp. This is not a property of the code; it is a
property of the storage layer. Data is stored with an
available_at field, and every query filters on
available_at <= as_of.
Vintage handling
Macroeconomic series are revised. GDP, employment, and inflation
series are subject to multiple revisions after first publication.
Using today's value of a series to backtest a decision made two years
ago is look-ahead, even if no future-dated observations are used.
Every series in the pipeline is graded:
| Vintage grade |
Meaning |
Eligibility |
| Exact |
Historical first-release values are preserved and retrievable for each publication date. |
Eligible for model features and backtesting. |
| Approximate |
Some vintage history exists but is incomplete, or has been reconstructed from revision logs. |
Eligible with a documented accuracy bound. Must be disclosed in the validation report. |
| Unavailable |
Only the latest revised value is retrievable. Historical vintages do not exist in any accessible source. |
Not eligible for model features used in backtesting. May be used for the live dashboard, with an explicit point-in-time caveat. |
Blocking issue. A model feature set is not considered
complete until every input has been assigned a vintage grade, and
every "unavailable" series has either been dropped or replaced. This is
currently the gating item for the regime model's first fit.
Data quality gates
Before a dataset is eligible for training, the following checks must
pass. Any failure halts the pipeline and produces a machine-readable
reason.
Schema conformity. Every series matches its expected schema: column names, types, frequency, and units.
Timestamp ordering. Observation, release, and availability timestamps are monotonic and consistent with the source's publication schedule.
Duplicate detection. No two rows claim the same series, observation date, and vintage.
Gap analysis. Missing observations within a series are enumerated and classified (weekend, holiday, delayed publication, data error).
Stale series detection. Series whose last observation is older than their expected refresh interval are flagged.
Future-date rejection. No observation may carry an available_at later than the dataset's stated availability cutoff.
Unit-change detection. Series that change units or frequency mid-history are flagged and require explicit handling.
Provider disagreement. Where the same conceptual series is available from multiple providers, values are compared and material differences flagged.
Section 04
Split design and leakage prevention
Time-series data cannot be randomly split. Adjacent observations are
correlated, targets overlap, and a randomly chosen test set contains
rows whose information was already visible to the model during
training.
Chronological splits
Every dataset is split chronologically. The default structure is
train–validation–test, with each split covering a contiguous block of
time and no overlap between them. Expanding-window and rolling-window
walk-forward are supported and preferred for final evaluation.
Purge and embargo
When the target is a forward return over horizon h, the last
h rows of the training set have targets that extend into the
validation set. Those rows must be purged. Immediately after each
split boundary, an embargo period is applied so that the first rows
of the validation set do not share information with the last rows of
the training set.
Training
→
Purge last h rows
→
Embargo gap
→
Validation
→
Purge last h rows
→
Embargo gap
→
Test
The test set is touched once
The final test set is not consulted during model development, feature
selection, hyperparameter tuning, or any other iterative process.
Consulting it more than once per artifact version converts it into a
validation set, and invalidates the out-of-sample claim.
Leakage tests
Every pipeline runs an automated leakage test suite. A test failure is
a hard block, not a warning.
Feature timestamp ≤ prediction timestamp for every row.
Feature available_at ≤ prediction timestamp for every row, including vintage dates.
Target horizon purge applied at every split boundary.
Scaler and normalizer fit on training only, then applied to validation and test.
Label encoders built from training categories only, unknown categories at inference treated as missing.
No target-derived feature enters the model. Rolling statistics computed on features must not include the target.
Shuffle test. Shuffle the target and re-fit. If the model improves on the shuffled target above chance, a leak exists.
Section 05
Baselines before models
Every predictive target is evaluated against a set of trivial
baselines before any non-trivial model is fit. A model that does not
beat its baselines out-of-sample has no claim on being a model.
Standard baselines
| Baseline |
Prediction |
Purpose |
| Historical mean |
Training-set mean of the target |
Establishes whether the model adds information beyond the unconditional expectation. |
| Historical median |
Training-set median of the target |
Same as above, more robust to outliers. |
| Persistence / zero |
Zero forward return |
The default for a strategy that does not trade. A model must beat not trading. |
| Most recent value |
Last observed target value |
Tests whether the model captures any momentum or mean reversion. |
The baseline comparison is not optional.
Validation reports list baseline metrics alongside model metrics, side
by side. A model that reports only its own metrics is not a completed
experiment.
Section 06
Validation gates
Every component must clear a set of gates before it can be promoted
between lifecycle stages. Gates are hard thresholds. Where a
threshold is not yet defined, the gate is listed as pending and no
promotion is possible until it is defined. The full gate set is
applicable to components that produce predictions; infrastructure
and retrieval components have a reduced gate set defined by their
class.
Data gates
| Gate |
Threshold |
Evidence required |
| Point-in-time correctness |
100% of rows pass available_at ≤ as_of |
Leakage test suite run log |
| Vintage coverage |
100% of features have a vintage grade assigned |
Feature registry with vintage column populated |
| Schema conformity |
Zero schema errors at ingest |
Data validation report |
| Sample size |
Threshold pending per component class |
Row count after all filters |
Statistical gates
| Gate |
Threshold |
Evidence required |
| Beats baselines out-of-sample |
Must exceed all four standard baselines on primary metric |
Side-by-side metrics table |
| Out-of-sample vs. in-sample gap |
OOS performance within a documented fraction of IS performance |
Overfitting diagnostic report |
| Multiple-testing adjustment |
p-values adjusted for the number of configurations tested |
Configuration count and adjustment method |
| Calibration |
Threshold pending per output type |
Reliability diagram, Brier score or equivalent |
Stability gates
| Gate |
Threshold |
Evidence required |
| Cluster or coefficient stability |
Stable under bootstrap resampling |
Bootstrap diagnostic, at least 200 resamples |
| Persistence over time |
Every defined state or cluster persists for a documented minimum |
Assignment timeline |
| Sub-period stability |
Performance not concentrated in a single historical regime |
Segmented performance report |
| Feature stability |
Coefficient signs and importances do not flip under resampling |
Feature diagnostics |
Coherence gates
| Gate |
Threshold |
Evidence required |
| Economic interpretability |
Every cluster centroid, coefficient, or state transition is describable in plain English |
Written interpretation accepted by review |
| Reference-set agreement |
Agreement with hand-labelled historical periods at a rate above chance |
Confusion matrix against the reference set |
| Dimensional consistency |
Every composite has a single semantic axis with all components aligned in sign |
Feature registry with sign convention documented |
Where a threshold is marked "pending", no promotion occurs.
A gate that has not been defined is not a gate. It is a placeholder.
Components whose validation reports contain pending gates remain in
the Research stage.
Section 07
Reproducibility and experiment tracking
Every experiment must be reproducible from an identifier plus the
recorded inputs. If the same experiment is run twice on the same
dataset, it must produce the same outputs. If it does not, the
discrepancy is a bug and must be resolved before the experiment is
treated as evidence.
What every experiment records
Experiment ID - stable identifier, unique across the registry.
Dataset ID - exact snapshot used, with row count and date range.
Feature version - identifier of the feature set used, with the versioned registry entry it resolves to.
Target version - identifier of the target construction, including horizon.
Model type and hyperparameters - full configuration, not just the class name.
Random seed - for every source of randomness, including initialisation and resampling.
Train, validation, test boundaries - including purge and embargo settings.
Code commit - git SHA of the repository at run time.
Dependency lockfile - exact versions of every library in the environment, not just the top-level names.
Environment - Python version, platform, and any other runtime-affecting detail.
Metrics - every computed metric, per split, in a machine-readable form.
Artifact hash - sha256 of the trained model, scaler, and feature schema, so the exact artifact is verifiable.
Artifact path - location of the trained model, scaler, and feature schema in the repository or model store.
Last retrain timestamp - when the current artifact was last produced, so staleness is visible.
Monitoring status - which drift signals are being tracked against this artifact, and their current state.
Status - current lifecycle stage and promotion history.
The reproducibility contract
An experiment is considered reproducible if, given the experiment ID
and the code commit, a full re-run produces:
The same dataset snapshot, with the same row count and date range.
The same feature matrix, element for element.
The same model configuration, including fitted parameters, to within the documented numerical tolerance.
The same out-of-sample metrics, to within the documented numerical tolerance.
The same artifact hash, or a documented reason why the hash differs (e.g. platform-level floating-point divergence).
The same validation report, modulo timestamps.
Reproducibility is a gate, not an aspiration.
A component whose experiments cannot be reproduced does not pass the
reproducibility gate and remains in Research regardless of its
metrics.
Section 08
Lifecycle and decision authority
Every component moves through a defined lifecycle. Promotion between
stages is a deliberate act, requires the evidence appropriate to the
target stage, and is recorded permanently in the registry. In
addition to lifecycle stage, every component carries a
decision authority - the maximum scope of action its
output is permitted to inform. Lifecycle and authority are separate
axes; a component can be in Production and still be scoped to
context-only authority.
Lifecycle stages
Design
→
Research
→
Candidate
→
Validated
→
Shadow
→
Production
→
Monitored
→
Degraded
→
Retired
| Stage |
Requirements to enter |
Consumed by |
| Design |
Specification complete. The question the component answers, its inputs, its outputs and its dependencies are documented. Implementation has not begun. |
Internal planning only. |
| Research |
Implementation in progress. Feature engineering, hypothesis testing, offline evaluation. No external output. |
Internal only. |
| Candidate |
Dataset, feature version, code commit, lockfile and configuration recorded. Artifacts committed to the repository. |
Internal only. |
| Validated |
All gates applicable to the component class passed. Validation report published internally. Human sign-off recorded. |
Internal research. Eligible for shadow deployment. |
| Shadow |
Minimum period of live-data operation with predictions compared to outcomes. No capital exposure, no live decision influence. |
Internal research. Publicly listed as shadow. |
| Production |
Shadow metrics pass defined thresholds. No regressions. Explicit sign-off. Decision authority assigned. |
Live use, scoped by decision authority. |
| Monitored |
Continuous production operation with all drift signals tracked, alerting active, and periodic re-evaluation scheduled. |
Live use, unchanged. Distinguished from Production by the completeness of monitoring. |
| Degraded |
Live performance departs from shadow expectations beyond documented tolerance, or a monitoring signal breaches its threshold. |
Output retained for analysis. Not used for decisions. |
| Retired |
Superseded, decommissioned, or found to be invalid. |
Historical reference only. |
No automatic promotion, ever.
The registry accepts promotion commands only from an authenticated
human action. There is no scheduled job, CI rule, or automated process
that moves an artifact between stages.
Decision authority
Every component carries a decision authority, independent of its
lifecycle stage. A component in Production with Context authority
cannot inform a position; a component in Shadow with Decision-support
authority is not yet permitted to act on it. Authority is assigned
explicitly and recorded in the registry.
| Authority |
Meaning |
Consumed by |
| Context |
Informs human reading. Makes no recommendation. |
Humans reading the dashboard. |
| Analytical |
Produces outputs consumed by other analytical components. |
Downstream components. |
| Forecast |
Produces forward-looking estimates with calibrated uncertainty. |
Downstream components and humans. |
| Decision-support |
Produces recommendations consumed by humans. |
Humans deciding whether to act. |
| Execution |
Authorised to affect positions or capital. |
The trading account. |
Lifecycle and authority are separate axes.
A component's stage says how mature its work is. Its authority says
what its output is permitted to inform. The two are recorded
independently in the registry so that a production component with
context-only authority is not mistaken for one authorised to drive
decisions.
Section 09
Shadow deployment
A component that has cleared its validation gates has proven itself on
historical data. It has not proven itself on live data. Shadow
deployment is the step between the two.
Evidence types — what "simulated" means
Five kinds of evidence appear at different stages of a component's
life, and they are not interchangeable. Where the registry or the
applied record describes a figure as simulated, it was produced by
one of the first three categories below — not by live trading.
Only the final category constitutes a live track record.
| Evidence type |
What it involves |
What it proves |
What it does not prove |
| In-sample backtest |
Strategy is fit and evaluated on the same historical data. |
That parameters exist which fit the given history. |
Nothing about future performance. Reported for completeness only. |
| Walk-forward validation |
Strategy is repeatedly fit on a training window and evaluated on the immediately following unseen window, rolled forward. |
That the edge, if it exists, is not a product of parameter selection on the tested period. |
That the edge will persist. Each window is an independent trial, not a live record. |
| Simulated holdout |
A single period is held back from all training and parameter selection, and evaluated once. The resulting figures are backtest output. |
That the strategy generalises to a period the model never saw, under stated execution assumptions. |
That the strategy has been deployed, that capital was at risk, or that real execution costs match the assumptions. |
| Shadow deployment |
The component runs against live data on schedule, producing outputs that are stored but not acted on. |
That the component operates correctly on live data, with outputs compared against real outcomes. |
That the component has influenced positions or capital. |
| Live track record |
The component or strategy is deployed and operating on real capital, with real execution. |
That the strategy performs under live conditions, with all associated frictions. |
That it will continue to perform. |
Where the applied record labels a period simulated holdout,
that label states which category of evidence the figures belong to.
It is not a claim of live trading. The registry uses this
classification when assigning lifecycle stages: a component that has
passed walk-forward and holdout evaluation but has not been deployed
is marked Validated, not Live.
What shadow means
A shadow component runs against live data on a fixed schedule. It
produces outputs on the same cadence as the production system
would. Those outputs are stored, timestamped, and preserved
alongside the input features that produced them. They are not used to
make decisions.
What is measured
Output quality
- Whether outputs match the eventual observed outcomes
- Whether calibration holds out-of-sample
- Whether uncertainty intervals contain the realised values at the expected rate
- Whether the distribution of outputs is consistent with the training-time distribution
Operational quality
- Whether the component runs on schedule without errors
- Whether its inputs are fresh and complete
- Whether its outputs are computable within the required latency
- Whether any data-quality warnings were raised during the shadow period
Duration and exit
Shadow duration is set by the component's output cadence and, where
applicable, the target horizon. A monthly regime classifier requires
a longer shadow period than a daily directional model because fewer
outcomes are observed per unit time. Exit criteria are defined before
shadow deployment begins; they are not adjusted mid-period to
accommodate the observed results.
Exit criteria are set in advance.
Adjusting shadow exit criteria after seeing shadow performance is a
form of overfitting. If the criteria need revision, the shadow period
restarts.
Section 10
Monitoring and drift
A component in production is not assumed to remain valid
indefinitely. The following are monitored continuously, and
departures trigger defined responses.
| Signal |
What is measured |
Response on breach |
| Data freshness |
Age of each input series relative to its expected refresh interval |
Flag stale inputs. Suppress component output if critical inputs are stale beyond tolerance. |
| Input distribution drift |
Population stability index or equivalent on each feature |
Record drift event. Escalate to research if sustained. |
| Output distribution drift |
Distribution of component outputs over a rolling window |
Compare against training-time distribution. Flag if materially different. |
| Calibration drift |
Reliability of probabilistic outputs against realised outcomes |
Recompute calibration on recent window. Downgrade if calibration fails. |
| Performance drift |
Outcome metric against expected value from validation |
Compare against shadow-period baseline. Trigger review if outside tolerance. |
| Operational failures |
Error rate, latency, retry count |
Standard operational response. Component output suppressed if failures exceed tolerance. |
Degradation is a defined state, not a failure.
A component whose live performance departs from its validation
expectations enters the Degraded stage. Degraded components continue
to produce outputs for analysis, but those outputs are not used for
decisions until the component is re-validated or retired.
Section 11
What is published, what is withheld
This page, and the system registry it governs, are intended to be
substantive. They are not exhaustive. The following boundary applies
to every public document.
Published
- The methodology itself: gates, splits, baselines, tracking practice
- The existence of a regime filter, in general terms
- The count and outcomes of walk-forward windows
- Forward-period summary statistics
- The identity of the canonical composites used as features
- The lifecycle stage and decision authority of each component
- Failures, corrections, and rejected approaches
Withheld
- Numeric values of any threshold, cutoff, or parameter
- Stop distances, entry bands, session blocks
- Position-sizing inputs and risk-budget constants
- Feature weights, scoring coefficients, sign conventions
- Exact dataset composition
- Internal artifact identifiers beyond registry entries
- Specific model hyperparameters
The boundary is deliberate.
Transparency about process and outcomes is compatible with discipline
about parameters. Both are required to run a research firm whose work
is both credible and defensible.
Section 12
Rejected approaches
A validation methodology is only as strong as the practices it
rejects. The following approaches have been considered and are not
permitted in any pipeline governed by this document.
Random train/test splits on time-series data. Adjacent observations are correlated; a random split leaks information from validation into training.
Fitting scalers on the full dataset. Scalers and normalizers must be fit on training data only and then applied unchanged to validation and test.
Reporting in-sample performance as evidence. A component's fit to data it was trained on is not a validation result.
Ignoring multiple-testing effects. Testing 50 configurations and reporting the best one without adjustment is a form of p-hacking.
Using revised macro data without vintage awareness. As described in Section 03, this invalidates any point-in-time claim.
Auto-promotion from validation to production. Every promotion is a deliberate, recorded human action.
Publishing validation claims before they are verified. A registry entry is a claim. Claims that have not been verified are not published.
Publishing operational history that did not occur. Deployment dates, evaluation dates and lifecycle transitions must reflect what actually happened. A status page that shows events on a timeline is not permitted to invent those events.
Adjusting thresholds after seeing results. Thresholds are set before validation begins. Changing them mid-validation invalidates the validation.
Treating one track's validation as evidence for another. Track A and Track B are independent. Neither validates the other.
Applying the wrong evidence type to a component class. A retrieval system is not validated by out-of-sample R². A statistical engine is not validated by walk-forward profit factor. Each class has its own evidence standard.
Silent corrections. When an entry is wrong, the correction is recorded, not silently overwritten.
Section 13
Governance
This document is a working standard. It is reviewed quarterly, and
any change to it is recorded in the system registry's changelog.
Ownership
This methodology is owned by the firm, not by any individual. The
founder is bound by it in the same way as any future team member. No
component is promoted without the evidence required by the sections
above, regardless of who is doing the work.
Challenging a result
Any validation claim may be challenged. A challenge is a request for
the raw data, the configuration, the code commit, the lockfile, and
the metrics that support the claim. If those cannot be produced, the
claim is withdrawn until they can. This applies to claims made in
public documents, in private reports, and in internal discussions.
Response to failure
When a validation claim turns out to be wrong, the response is:
Record the correction in the changelog, dated and attributed.
Determine whether the error was in the evidence, the interpretation, or the process.
If the process is the source of the error, amend this document.
Re-validate any component that relied on the withdrawn claim.
A methodology is only credible if it is applied when the
results are inconvenient.
Rules that are dropped when they produce bad outcomes are not rules.
This document exists so that the temptation to make an exception is
visible when it arises.