Back to Hylios

Methodology

How Hylios validates a simulated supply chain

Hylios builds a working model of a company's supply chain from public data, without that company's cooperation. Anyone evaluating a number we produce is entitled to ask two questions: where did it come from, and what was checked before you showed it to me. This page answers both, including the parts that are unflattering.

Hylios ResearchPublished September 9, 2026

The entire product rests on the claim that our numbers are generated, checked and auditable. A methodology page that oversells would destroy the thing it is meant to support, so the limitations below are stated as plainly as the capabilities.

The three kinds of number in any Hylios output

Every figure we show is one of these. They carry different weight and should be read differently.

We do not present estimated figures as observed ones. Every model call that produces a number is logged with a hash of its prompt and its response, so any specific figure can be traced back to what produced it.

What gets checked

Every completed run is graded automatically by a validator that currently runs 66 distinct checks. It runs after metrics are computed, and its verdict is stored with the run rather than recomputed later. Nobody has to remember to run it.

It grades both the baseline and the what-if scenarios. That is worth stating explicitly because it was not always true: until recently only the baseline was sent for grading, which meant every comparison between a baseline and an alternative had one graded side and one ungraded side. If we show you a what-if, it has been checked.

  • Structural. Does the modelled network actually cover the company: are suppliers geocoded, do manufacturing nodes have inputs, does every network role have catalog data behind it, is the supply chain as broad as the company's real one.
  • Conservation. Does volume balance. Demand that enters must be served or explicitly booked as unserved. Flow into a stage must leave it, except where transformation is physically expected. Demand that simply disappears is a defect, not a saving, because it silently flatters every per-unit figure computed over it.
  • Economic plausibility. Cost per unit, cost against sales price, the split between transport, holding and handling, mode cost ordering (air should not be cheaper than ocean), emissions intensity, tariff upper bounds, lead times.
  • Engine self-checks. The simulation's own invariants: capacity never exceeded, flow conserved, demand satisfied. These run inside the engine while the numbers are still in memory, and their results are carried out to the verdict.
  • Coverage and provenance. Whether the checks themselves ran, how much of the catalog the demand footprint covers, whether the evidence behind the run is present.

Verdicts

A run is graded PASS, PASS-WITH-NOTES, or REVIEW.

  • PASS: no check failed and no advisory note was raised.
  • PASS-WITH-NOTES: the numbers held up, but the run carries caveats worth reading.
  • REVIEW: at least one check failed, or the run could not be graded. It needs a human before it is shown to anyone.

The mapping fails closed: any verdict state the system does not recognise reads as REVIEW, never as PASS. A new or unexpected internal state cannot be mistaken for a clean run.

A check that did not run is not a check that passed

A check that did not execute is reported as not having executed, rather than counted as a check that found nothing. Those are different statements, and conflating them is how a broken pipeline looks healthy.

That principle has cost us real work, which is the only evidence that we mean it. When the high-volume detail behind several checks was moved out of the database into object storage for speed, fourteen checks were left reading tables that were now empty. Each returned "nothing found" instantly, and from the outside that is indistinguishable from passing. All fourteen now read the moved data. When a check is skipped because reading its evidence would exceed the time a run is allowed, the run says so by name, and the verdict is presented as resting on fewer checks than usual.

What we will not claim

  • Not audit grade. Nothing here is prepared to an assurance standard. Emissions output in particular is screening-grade Scope 3 Category 4 estimation from public data, and must not be presented as an audited inventory.
  • The twin is not the company. It is a model built without the company's cooperation, from public information. It is useful because it is directionally right and internally consistent, not because it is a copy.
  • A passing verdict does not mean the answer is correct. It means the specific checks we run did not find a problem. Checks find defects; they cannot certify a model.
  • No precision we cannot support. Where a figure is estimated, it is labelled estimated. Where a range is honest, we give a range.
  • We do not hide a REVIEW. A run that fails validation is not quietly reissued as a clean one.

Known limitations, stated plainly

These are current as of 2026-09-09 and will change. That is the point of dating the page.

  • No run has yet been calibrated against a company's actual figures. Establishing that, and publishing the deltas including the misses, is a separate piece of work. Until it exists, we can tell you what we check, not how close we get.
  • Tariff attribution is being rebuilt, and we can say how badly. The engine has taxed geography rather than the importer of record. The rate half is fixed; incidence is not. In a sample of fourteen graded scenarios, three carried genuinely dutiable flow with a total tariff cost of exactly zero: in one case 3.3 million units moving Vietnam to the United States, a lane pair with no free-trade agreement to explain the zero. Duty is being under-attributed. Read tariff figures as indicative until this completes.
  • Product classification is weak. Matching a product to a tariff code is measured at 22.8% top-1 accuracy in the current shipping path, with a two-stage replacement measured at 72.8%. Measured, because similarity scores were being mistaken for accuracy for a long time.
  • Demand footprints are narrow. Completed demand maps have been observed covering between 0.8% and 4.7% of the catalog considered, and two runs with identical persisted parameters produced eligible sets differing by six times. Both are open defects.
  • Constraints the customer sets are not always honoured. Where a run carries a maximum lane utilisation, we have measured up to 77% of lane-weeks exceeding it. The flow allocator treats it as a preference rather than a limit. This is now checked on every run, so it appears in the findings instead of going unnoticed.
  • The validator is stricter than the engine is mature. Across every graded scenario in our development environment, no run has ever earned a PASS. We would rather publish that number than a verdict distribution we curated. It also means the gate that would withhold a failing run from being shared is deliberately running in advisory mode, recording what it would have blocked rather than blocking it, because a gate that stops everything on day one teaches people to route around it.

How to challenge a number

Ask for the run's verdict and its findings. Every graded run stores them. Ask which of the three kinds a figure is, and if it is estimated, ask what produced it: the model audit log records prompt and response hashes for every generated value.

If a number cannot be traced, treat it as unsupported. That is the standard we hold ourselves to, and it is the only one worth publishing.