Ant Says Its FX Model Tops the Leaderboard. We Downloaded the Leaderboard.

A company claimed state-of-the-art performance on a public leaderboard and did not say which leaderboard. The raw results of the one it means are downloadable by anyone. We downloaded them and recomputed the rankings, and the fuller picture turns out to have been published by the company itself six days earlier.

Last updated: August 21

Key takeaways

Link copied
  • Ant International says FalconTST 2.0 tops a leaderboard on MASE. It does not name the leaderboard; its GitHub does.
  • Recomputed from GIFT-Eval's raw results, Falcon-2.0 scores 0.6660 MASE — first among the 35 pretrained entries.
  • On the same benchmark's other headline metric, CRPS, the model places twelfth of those 35 entries.
  • Across all 122 scored entries rather than one category, the same MASE score ranks tenth.
  • Ant's own technical report is more careful than its press release and reports the weaker results itself.
  • The release names four banks. Coverage reporting six banks, HSBC, or a 60% saving is unsourced in Ant's materials.

Data highlight

10rank out of 122 scored entries

Rank of Ant International's Falcon-2.0 on GIFT-Eval normalised MASE across all scored leaderboard entries

as at 2026-08-21

Download of every per-configuration result file published in the GIFT-Eval leaderboard repository, followed by independent re-aggregation: each model's MASE is divided by the seasonal-naive baseline for the same configuration and the ratios are combined as a geometric mean across all 97 dataset-frequency-horizon configurations, matching the benchmark's stated method. Falcon-2.0 records a normalised MASE of 0.6660 — first among the 35 entries labelled pretrained, but tenth of the 122 entries that carried enough data to score across all six categories. Nine entries score lower: eight labelled agentic and one fine-tuned, led by EXAONE-Forecast-Agent at 0.6099. One of the nine, Falcon-Agent at 0.6658, was submitted by Ant International itself. On CRPS, Falcon-2.0 records 0.4863 and places twelfth among the 35 pretrained entries. Category labels are recorded by each submitter and are not audited by the benchmark. Method check: this recomputation reproduces a CRPS of 0.4544 for STRIDE (+Chronos-2), the same figure printed in Ant's own technical report.

For us, the value of AI is not simply achieving a better forecasting score, but turning that predictive intelligence into real decisions—how much liquidity to prepare, how to manage FX exposure, and how to allocate capital more efficiently.
Jiang-Ming Yang, Chief Innovation Officer, Ant International

Ant International announced on 19 August that its FalconTST 2.0 forecasting model had achieved state-of-the-art performance, and that four global banks had integrated it into their foreign exchange operations. The announcement said the model "achieved a MASE score of 0.666 and places it at the top of the leaderboard, surpassing other TST foundational models from leading global tech companies."

It did not say which leaderboard. We found it, downloaded the raw results, and recomputed the rankings.

The benchmark the announcement does not name

Neither the press release nor Ant's FalconTST product site names the benchmark. The product site says only "1st Rank in Zero-Shot on Benchmark."

Ant's own code repository names it plainly. The README for Falcon-TST on GitHub carries the line "Comparison of MASE on the GIFT-Eval benchmark." GIFT-Eval is a public evaluation suite published by Salesforce AI Research, covering 23 datasets across seven domains and ten sampling frequencies, which combine into 97 dataset–frequency–horizon configurations. It scores two headline metrics: MASE for point forecasts and CRPS for probabilistic forecasts, each normalised against a seasonal-naive baseline and geometrically aggregated.

Every model's per-configuration results are published as a file anyone can download. We retrieved all of them on 21 August 2026 — 123 entries — and re-aggregated them ourselves. As a check on the method, our recomputation reproduces a CRPS of 0.4544 for STRIDE (+Chronos-2), exactly the figure printed in Ant's own technical report.

The headline claim holds

Falcon-2.0 records a normalised MASE of 0.6660. That is the lowest of the 35 entries labelled pretrained, ahead of STRIDE (+Timer-S1) at 0.6744 — a margin of about 1.3%. The 0.666 in the press release is the current leaderboard figure, and on that metric, in that category, the model is first.

It is worth noting that this is a live number. Ant's technical report, posted six days before the announcement, gives 0.6684 and compares against 29 pretrained models as of July 2026. There are now 35. Both the score and the field have moved.

Figure 1. GIFT-Eval, pretrained category, recomputed by HaiPay from the benchmark's published raw results on 21 August 2026. Falcon-2.0 is first on MASE and twelfth on CRPS among the same 35 entries. Source: Salesforce GIFT-Eval. Chart: HaiPay.


The other metric on the same benchmark

GIFT-Eval reports CRPS alongside MASE. On CRPS, among the same 35 pretrained entries, Falcon-2.0 places twelfth, at 0.4863. STRIDE (+Chronos-2) leads at 0.4544, and eleven models score better than Falcon-2.0, including three sizes of Toto-2.0, Timer-s1, chronos-2 and Ant's own Falcon-X.

Ant's technical report does not hide this. It states that Falcon-2.0 secured "the seventh-lowest CRPS (0.4843)" and that "STRIDE + Chronos-2 retains an edge on this specific dimension." The report was describing the July field of 29; on today's field of 35, seventh has become twelfth.

Tenth across the whole leaderboard

The pretrained label is one of six on GIFT-Eval. The others are zero-shot (39 entries), agentic (27), deep-learning (10), fine-tuned (6) and statistical (5). The categories describe what a submission is: a single pretrained foundation model, an ensemble or agent that orchestrates several, a fine-tune, a classical statistical method.

Ranked against every scored entry rather than one category, Falcon-2.0's 0.6660 places tenth of 122. Nine entries score lower on the same metric — eight agentic systems and one fine-tune. The leader, EXAONE-Forecast-Agent, records 0.6099, about 8.4% better.

One of the nine is Ant's. Falcon-Agent, submitted by the same organisation in the agentic category, records 0.6658 — fractionally ahead of the model the announcement is about.

None of this makes "top of the leaderboard" false. It makes it a statement about a category, and the announcement does not say which one.

Figure 2. The twelve best normalised-MASE scores across all 122 scored GIFT-Eval entries. Falcon-2.0 leads its own category and ranks tenth overall; Ant's own Falcon-Agent scores marginally better. Source: Salesforce GIFT-Eval, recomputed by HaiPay. Chart: HaiPay.



The technical report is more careful than the press release

The gap between the two documents is instructive. The paper, "Into the ORBIT for Time Series," was posted to arXiv on 13 August. Its abstract claims "strong zero-shot forecasting performance." It does not use the phrase state of the art about the model's overall standing; the one appearance of the term is narrowly scoped to probabilistic calibration on a second benchmark.

On that second benchmark, fev-bench, the paper reports Falcon-2.0's aggregate MASE of 0.6459 as "within 0.3% of the top-performing TimesFM-2.5 (0.6438)" — that is, second. It also reports that on the 46 tasks that carry known-future covariates, Chronos-2 beats Falcon-2.0 on both metrics, because Falcon-2.0's interface does not ingest future features. The paper is explicit that the residual gap "stems primarily from covariate conditioning limitations."

Everything a sceptical reader would want is in Ant's own paper. None of it is in the announcement.

Why the metric split matters for hedging

The distinction between the two metrics is not a technicality for the stated use case. MASE measures how close a single predicted number lands. CRPS measures how well a predicted distribution is calibrated — whether the range the model gives you is honest about its own uncertainty.

Hedging is a distributional decision. The press release makes the point itself: a company that forecasts too high over-hedges, and one that forecasts too low leaves exposure open. Sizing that trade-off requires knowing the spread of plausible outcomes, not just the midpoint. The metric on which FalconTST 2.0 leads is the one about midpoints.

In fairness, the picture is benchmark-dependent, and the paper says so: on fev-bench, Falcon-2.0 records the best aggregate weighted quantile loss of any model tested while completing all 100 tasks. Its probabilistic performance is strong on one suite and mid-field on the other.

Four banks, and the numbers that grew in the retelling

The announcement names four banks: Barclays, which uses the model in its BARX NetFX platform; Citi, which pairs it with its Fixed FX Rates product; Deutsche Bank; and Standard Chartered, which runs it alongside its SCALE FX system as part of both firms' participation in the Monetary Authority of Singapore's PathFin.ai programme.

Ant's own product site lists four success stories, dated between May and October 2025: Barclays, Citi, Standard Chartered and the airline group Capital A. Deutsche Bank is not among them.

Secondary coverage has reported six banks including HSBC, and a claim that the forecasts cut FX hedging and allocation costs by more than 60%. Neither the six-bank count, nor HSBC, nor the 60% figure appears in the release or on the product site, and we could not source them. HSBC does have a documented relationship with Ant International, on a tokenised deposit service that processed a cross-border ISO 20022 payment in 2025 — a different programme. The two appear to have been merged somewhere in the retelling.

The accuracy figure moves too. The release says forecast accuracy above 93%; the product site says above 90%. Neither defines what is being measured, over what horizon, or on whose data.

What is established and what is not

Established: Ant International announced FalconTST 2.0 on 19 August 2026. The benchmark behind the claim is GIFT-Eval, named in Ant's GitHub repository. Recomputed from the benchmark's published raw results on 21 August 2026, Falcon-2.0 records a normalised MASE of 0.6660, first among 35 pretrained entries and tenth among 122 scored entries overall; and a normalised CRPS of 0.4863, twelfth among the pretrained entries. Ant's technical report states a MASE of 0.6684 and a seventh-place CRPS against the July field, and reports being second to TimesFM-2.5 on fev-bench point accuracy. The release names four banks.

Not established: what the ">93% forecast accuracy" figure measures; the basis for the "over 60%" cost-reduction claim, which we did not find in Ant's own materials; whether any bank beyond the four named has deployed the model; and how any of the benchmark results translate to accuracy on live cross-border FX flows, which no public benchmark tests. Our category and ranking figures depend on labels each submitter recorded for its own entry, which the benchmark does not audit.

The narrow reading is that a company made a checkable claim about a public leaderboard, the claim survives checking within the category it silently refers to, and the fuller picture — which is less flattering and more interesting — was published by the company itself, six days earlier, in a paper the announcement does not mention.

How to cite

Link copied

HaiPay News, "Ant Says Its FX Model Tops the Leaderboard. We Downloaded the Leaderboard.", https://www.haipay.net/news/falcon-tst-2-gift-eval-benchmark-recomputed, August 21st, 2026

About the author

Crystal

Digital Public Relations

A digital PR specialist with a Master's in Journalism & Communication from UNSW. Started as an intern at ABC Australia, now leads public relations at Haipay, crafting press releases and media strategies that bring brand stories to life.

Reviewed by WeiJun TangEditorial policy

6 sources

Discover More