applications / firms
an application · a new population
The firms.
Microcosm builds weighted synthetic populations where policy-relevant microdata is missing. Households were the first population; this is the second. The UK's VAT registration threshold is a cliff in the firm-size distribution, published statistics resolve it only in coarse turnover bands, and the firm-level data needed to cost reforms is not available in open form — so the same recipe applies: draw records from the published structure, calibrate weights to administrative aggregates, simulate reforms on the file.
Three reform families, one common base.
Once turnover crosses £85,000, VAT falls due on a firm's whole base — a notch, not a kink. Reform proposals change the threshold's level (move it), its shape (taper the rate in), or its rate (a reduced-rate band above it), and the three are not comparable through the reduced-form methods used to study the threshold. A firm-level file makes them comparable by construction: every reform is costed statically on the same 2,407,267 weighted firms and the same £175.4bn 2023–24 registered base.
firm-microsim-paper, 2023–24 vintage; seed half-ranges ±£1.4m (raise) and ±£0.3m (taper) across generator seeds. The taper's wide band is forced, not chosen: removing the dominated region while confining relief to the band pins the band top at £141,667 — a taper reaching the full rate by £105,000 recreates a dominated interval smoothly. Conditional on an assumed turnover elasticity, intensive-margin re-optimisation barely moves these figures in the paper's region-confined model: a threshold raise there has zero intensive-margin offset — released firms leave the VAT base, so their expansion is untaxed — and the reduced-rate bands recover at most £18m of £515m at the highest elasticity swept. The static menu is the costing.
The household recipe, on business registers.
The pipeline is two stages, parameterised by a single threshold. Draw: sample 2,941,275 firm records — continuous within-band turnover, employment, sector, and intermediate inputs — from the ONS business-register structure, so the file has firm-level resolution the published bands lack. Calibrate: fit the weights by the same optimisation family as the household stack (Adam on a symmetric relative-error loss) so weighted totals reproduce HMRC VAT-registered counts by turnover band and sector, ONS employment bands, HMRC liability totals, and the OBR's £1,000-band counts near the threshold. Registration is mandatory above the threshold and voluntary below at the HMRC-calibrated rate.
Each firm carries a structural net liability — the 20% rate applied to value added, τ(turnover − inputs) — rather than a residual the optimiser is free to absorb. An earlier build set liability to value added itself, five times too large per firm; aggregate calibration scores barely moved while the near-threshold density distorted. The lesson is a design rule the household stack shares: build liabilities from structured components, and never let the weight optimiser determine what an accounting identity should.
results/calibration_accuracy.txt; Kleven–Waseem dominated-region width at τ=0.20. Because the population is calibrated to the HMRC aggregates, agreement with them is an internal consistency check, not external validation. The frame is the ONS register of VAT- and/or PAYE-registered businesses (~2.7M): unregistered sole traders below the threshold are out of frame, so cross-threshold level comparisons against all-business tabulations will differ by construction.
Band-calibrated data carry no bunching evidence.
The generator has no location-choice model, so any excess mass below the threshold is inherited from the calibration targets, not discovered. The paper reports what the estimator reads back as an audit of the construction, not as evidence: on the 2023–24 build the baseline specification returns E = 10,418 firms (bLLAT = 5.13) — a readback of the OBR £1,000-band profile the population is calibrated to. Across the paper's degree-and-window grid the same statistic swings from zero to about 25,100, and the fitted missing mass above the threshold is zero in the headline specification. The readback is not a second estimate of the administrative bunching ratio (b = 1.361 in Liu, Lockwood, Almunia and Tam's HMRC micro-data); the paper therefore computes no elasticities from bunching.
results/bunching_inference.txt, seed_sensitivity.txt. On the 2024–25 vintage the estimator reads E = 25,227 at the £90,000 threshold — the same target-inheritance object on the newer vintage. No standard errors are reported because there is no estimand: the quantity is a property of the calibrated file, not of firm behaviour.
From standalone paper to Microcosm population.
The paper ships as a standalone repository — generator, calibration, estimator, and manuscript, regenerable end to end. Its official target surface is being mirrored into PolicyEngine's Ledger fact store, and an experimental UK firm generator already sits in Microcosm behind the household stack's gates; the repository includes a comparison command so the pinned migration snapshot can be audited against the published paper population. microcosm#258 tracks porting the corrected liability identity before any Microcosm firm build is certified, and an adversarial referee review is open on the paper repository (firm-microsim-paper#32).