85K-parameter time series foundation model

Fracast-0: Fractal Weight Sharing for a Time Series Foundation Model with Only 85K Parameters

Small enough to load directly in a web page.

Tianxiang Zhan1, Huanyao Zhang2, Yuanpeng He†,2

1 University of Electronic Science and Technology of China 2 Peking University † Corresponding author

“We built this solely to explore whether model compression can be pushed to an even more extreme state. We spent 15 days on this exploration. Although the work is not perfect, it is at least usable, so we are releasing it. Fracast-0 is the published version, and future versions will only get better.”
— Tianxiang Zhan

The idea in one page

01

Temporal structure repeats

A daily cycle, a weekly shape, and a slower trend can express similar local structure at different resolutions. Fracast-0 treats that self-similarity as a compression opportunity.

02

Reuse one operator

Instead of giving every temporal scale a separate set of weights, the model applies one shared causal dilated block along a geometric dilation ladder. FiLM conditioning tells the block which scale it is operating on.

03

Keep probability

The decoder combines context-gathered states with an explicit seasonal future state and emits nine quantile forecasts from 0.1 through 0.9.

Model architecture

Diagram showing Fracast-0 preprocessing, shared-scale encoder, and decoder producing quantile forecasts
Fracast-0 turns a univariate context into nine quantiles in three stages: preprocessing with a parameter-free seasonal detector, an encoder with a shared block over a dilation ladder, and a decoder with an explicit future-state path.

Preprocess

Normalize each window, preserve missingness, and derive bounded recency features plus a seasonal prior.

Encode

Project seven input features to the hidden width, then reuse the same local block across scales with scale conditioning.

Decode

Gather relevant context states, combine them with a periodic future state, and map the result to the target horizon.

How well does it work?

85,001Parameters
−42.0%Parameters versus TinyCast
0.807Overall MASE
0.563Overall WQL
23.9 msMedian CPU latency
97GIFT-Eval configurations

Scores use the official GIFT-Eval protocol without per-dataset fine-tuning. The model is pretrained on corpora that overlap GIFT-Eval families, so these results are labeled pretrained rather than strict zero-shot. Lower MASE and WQL are better.

Fracast-0 results across GIFT-Eval, TIME, FEV-Bench, and BOOM
Benchmark Headline result Scope
GIFT-Eval Normalized MASE 0.807133 Normalized MWQL 0.563008 97 configurations; official results directory
TIME Normalized MASE 0.767965 Normalized CRPS 0.649192 98 tasks; rank 24/29 on both metrics in the official table
FEV-Bench Controlled MASE rank 16/30 Controlled SQL rank 16/30 100 tasks; raw ranks MASE 20/30 and SQL 17/30; in-corpus
BOOM Scaled MASE 0.723 Scaled CRPS 0.434 7,413 configurations; in-corpus

Scores use the official evaluation protocols and the official comparison tables for each benchmark. FEV-Bench and BOOM are in-corpus because the pretraining recipe includes their evaluation datasets. TIME uses median-quantile feedback beyond 48 steps on 47 of 98 tasks.

Scenario highlights

Error ratios compare each configuration or task with Seasonal Naive. Lower is better, and TIME ranks use the official 29-model table.

  • GIFT-Eval, Sales: All four short-horizon Sales configurations improve on Seasonal Naive in normalized MASE and normalized MWQL. The geometric means are 0.693 and 0.422, and Fracast-0 leads TinyCast on both metrics in every configuration.
  • GIFT-Eval, Cloud operations: Across the three hourly bizitobs_l2c short, medium, and long slices, normalized MASE is 0.430 and normalized MWQL is 0.353.
  • TIME, Solar forecasting: Australia_Solar/H ranks 4th to 6th of 29 by MASE and 7th to 13th by CRPS over short, medium, and long horizons.
  • TIME, Manufacturing: On Smart_Manufacturing/H, medium and long horizons rank 8th of 29 by MASE, while CRPS ranks range from 8th to 10th across all three horizons.

These are scenario slices, not separate aggregate rankings. The benchmark disclosures above still apply.

Scatter plot of normalized MASE against parameter count
Normalized MASE across released checkpoints. Lines connect checkpoints from the same model family.
Scatter plot of normalized weighted quantile loss against parameter count
The same comparison on WQL. Fracast-0 occupies the extreme low-parameter end.
Parameter-accuracy Pareto frontier plots
Fracast-0 is non-dominated among 28 evaluated checkpoints. Moving from Fracast-0 to TinyCast spends 1.72 times more parameters for 4.2% lower MASE and 3.3% lower WQL.
Plots of normalized MASE versus parameters for short, medium, and long horizons
Short-horizon performance is close to TinyCast. Medium and long horizons remain the main room for improvement.
Bar chart of batch-one CPU latency
On a fixed local CPU workload, Fracast-0 reports 23.9 ms versus 41.4 ms for TinyCast.
Bar chart of resident CPU memory
The same workload reports 389 MiB versus 471 MiB resident memory.

Qualitative forecasts

Six selected GIFT-Eval windows cover smooth levels, abrupt load changes, daily weather cycles, periodic solar generation, and noisy event counts. The shaded band is Fracast-0’s 10–90% interval.

Legend for qualitative forecast comparisons
Qualitative short-horizon forecast panel one
Qualitative short-horizon forecast panel two
Qualitative short-horizon forecast panel three
Qualitative short-horizon forecast panel four
Qualitative short-horizon forecast panel five
Qualitative short-horizon forecast panel six

Try it locally

python -m pip install fracast

Good fit

Univariate or channel-independent forecasting, memory-constrained deployment, probabilistic outputs, and short-horizon scoring on CPU.

Current limits

The released model is channel independent, natively predicts 48 steps per block, and needs rollout for longer horizons. The W8 export dequantizes weights and is not an integer-only kernel.

Just try the web demo

Run Fracast-0 in your browser with preset traffic data. No installation or GPU required.

Open the web demo

Citation

@article{zhan2026fracast0fractalweightsharing,
  title={Fracast-0: Fractal Weight Sharing for a Time Series Foundation Model with Only 85K Parameters},
  author={Tianxiang Zhan and Huanyao Zhang and Yuanpeng He},
  year={2026},
  journal={arXiv preprint arXiv:2609.32209},
  eprint={2609.32209},
  archivePrefix={arXiv},
  primaryClass={cs.AI},
  doi={10.48550/arXiv.2609.32209},
  url={https://arxiv.org/abs/2609.32209}
}

References and resources

Thanks

We thank Xiaomi MiMo V2.6 Pro and DeepSeek V4.1 Flash for supporting the implementation, debugging, and release work behind Fracast-0.

Sponsorship

We are seeking sponsors for continued research, API credits, and server resources. Contact us at .