Cycle 4 submissions failing with "Failed to run" — reproduced clean locally
I've had two submissions to Cycle 4 fail with Failed to run:
submit4.py— 02/10/2026 18:10 UTC — 1.35 KBsubmit5.py— 02/10/2026 18:19 UTC — 2.35 KB
Both run cleanly on my machine with the shipped local kit:
Per-period Sharpe: 31/31 positive, ranging +0.026 to +0.129.
--gauge-fix also completes: IC=0.0641, concentration=0.0700, city novelty=36.3°.
Runs in ~0.2s per period on a Windows laptop with 8 GB RAM, Python 3.13.
What my Predictor does:
train(features, target): readstarget.iloc[-4:], stores the column-wise mean inself.cached, savestarget.columnsinself.ticker_order.predict(features): builds a constant prediction fromself.cached, multiplies by a linear decay factormax(0, 1 - t/200), cross-sectionally de-means, returns.
No top-level code except predictor = DecayCachePredictor(). No external dependencies beyond numpy and pandas. No shift(-n), no bfill, no centered windows. predict() is deterministic and handles any frame length via np.arange(n).
What I've already tried to fix it:
- Removed an
assert self.trainedfrompredict()in case the pipeline reloads from a pickled state. - Added NaN/inf handling in
train()(np.nanmean+np.nan_to_num) in case the shuffled-target gate produces all-NaN columns. - Added a length guard in
predict()in case the frame is padded.
Both failed versions produced the same Failed to run error.
Two questions:
- Is the Cycle 4 submission pipeline currently having issues, or is there a specific check I can run that catches what's failing server-side?
- If it's my code, is there a way to see a stack trace from the pipeline? "Failed to run" is a catch-all and I've hit every case it lists except a genuine pipeline error.
Thank you.
1 Reply
Hi @Mr_9oob, thanks for the detailed write-up – it made this quick to pin down.
The pipeline is fine: other participants' submissions were scored normally today. Both of your runs (submit4.py and submit5.py) failed for the same reason, and it isn't an exception in your code – your code ran, but its predictions were exactly zero on every scored hour of every month. An all-zero signal has nothing to score, and that case currently surfaces as "Failed to run".
Why the server sees zeros and your machine does not: in scoring, predict() is called once per month on that month's whole feature frame – the full history plus the month, roughly 32K–62K rows – and only the month's final rows are scored. Your decay is keyed on the row number within that frame (idx = np.arange(len(features)), decay = max(0, 1 - idx/200)), so it is 0 for every row after the 200th, which includes every scored hour. The local kit's walkforward.py passes only the validation slice to predict(), so there idx starts at 0 and the first 200 rows carry your bet – that's where the 31/31 positive periods come from.
Two rules to keep in mind before you resubmit (see the rejection reasons on the competition's Overview tab):
- ROW_MASK_DETECTED – predictions must not depend on where a row sits in the frame. A model scaled by row number or by a count of rows since the start of the frame is rejected even when its output is non-zero; anchor anything you carry forward to timestamps and a finite look-back.
- FLAT_PERIODS – a forecast that sits at zero for long stretches (e.g. a decay that reaches 0 and stays there for the rest of the month) is rejected as well.
You're right that "Failed to run" is too coarse for this case – we'll look at reporting it with a specific reason.