exploit cheating??
Evidence: exploit / konstik twin-account operation (Q1 holder)
Finding
The Q1 quality signal (exploit / signal_1355decf, sharpe 0.0522) is one of
nine bit-identical signal pairs submitted across two accounts, exploit
and konstik (10 signals each, 20 total — double the 10-per-scientist slot
budget).
| exploit signal | konstik twin | mutual-nearest | Δ ic | Δ compression_loss | Δ sharpe | sharpe |
|---|---|---|---|---|---|---|
| signal_1355decf | signal_87692c46 | yes | 0.00e+00 | 1.9e-15 | 1.3e-05 | 0.0522 (Q1) |
| signal_6b1ef723 | signal_f75b3a3a | yes | 0.00e+00 | 3.7e-15 | 2.7e-05 | 0.0505 |
| signal_fe07f3c7 | signal_9b9a02c6 | yes | 0.00e+00 | 2.8e-16 | 2.0e-05 | 0.0496 |
| signal_35b1fbf4 | signal_188e712c | yes | 0.00e+00 | 7.3e-15 | 3.6e-05 | 0.0494 |
| signal_8d1f9d89 | signal_0877f8b1 | yes | 0.00e+00 | 1.7e-15 | 3.8e-05 | 0.0492 |
| signal_45d19261 | signal_1fb804ad | yes | 0.00e+00 | 3.8e-15 | 1.4e-05 | 0.0484 |
| signal_f22bb2b6 | signal_626bb4b7 | yes | 0.00e+00 | 1.4e-15 | 1.9e-05 | 0.0473 |
| signal_b9ad78f0 | signal_04468804 | yes | 0.00e+00 | 1.7e-15 | 2.6e-05 | 0.0463 |
| signal_f694e426 | signal_c24ddb8c | yes | 0.00e+00 | 2.2e-16 | 1.8e-05 | 0.0428 |
- Δ ic is EXACTLY zero in all nine pairs; Δ compression_loss is floating-point epsilon. These are identical prediction vectors, not similar models.
- Each pair reports the other as
nearest_signal_publicANDnearest_city_publicat 0.000000° (this is what makes Q1 display "city novelty 0" — the "city" is the twin). - The tiny Δ sharpe (~1e-5) is consistent with the two copies being scored a day apart (different n_periods), i.e. two separate submissions of one file.
- exploit's 10th signal (signal_af5520d1) is 11.9° from ANOTHER exploit signal; konstik's 10th is 30.3° novelty. Both accounts' portfolios are one iterated family.
Why it matters
- Two accounts = 20 submission slots vs the 10 everyone else gets — doubled iterative probing of the (partially hidden) scoring key. exploit holds 7+ of the top raw scores 0.046-0.052; the density of that ladder is consistent with slot-doubling, not a distinct alpha.
- Multi-account / sockpuppet entry is a written DQ offense in the comp-5 rules (see tmp/comp5_rules.html.txt). Identical submissions across two "different" scientists is either one person with two accounts or two people sharing signal files — both prohibited.
- Consequence if enforced: exploit's Q1 (0.0522) is removed; the legitimate quality ladder is TaroKogawa 0.0460 / hctredfield37 0.0439 / roguetrader 0.0417 / radiant-allomancer 0.0411.
Conclusion for our own decision
The Q1 anomaly is NOT an undiscovered constructible edge in the features — it is a rules-violation pattern. This is consistent with all local findings (construction search closed at ~0.051 in-regime crowded-robust).
12 Replies
I had the same suspicion
To the competition admin : Is there a daily limit system ( like 5 submissions a day ) possible like kaggle ? because 10 is still to less in our whole experimentation flow .
Hi braxtonmensah, TaroKogawa and ayushkhaire,
thanks for flagging, we will take a look and revert! by the way, this is an example of why I think we need a DAO token to distribute.
Best
Marc
Again, thanks for the detailed report — this is exactly the kind of community vigilance the competition benefits from.
We want everyone to know this was taken seriously and acted on immediately. We have completed a thorough independent audit, re-deriving everything from primary sources (submission records, scoring artifacts, and account data) rather than relying on the post's analysis. The audit substantiates the core technical finding: the flagged signal pairs across the two accounts are identical submissions.
However, the audit also surfaced facts the post does not capture — including the precise timeline of events — that bear directly on what a fair and proportionate outcome looks like. We're not going to adjudicate this in the forum, and we'd ask everyone to hold off on conclusions until the process finishes.
Per the Competition Rules, we have reached out to both account holders with a formal request for clarification, on a short deadline. Once that window closes, we will make a decision under the rules (§6.A Individual Account, §8.A Private Code Sharing) and post the outcome and reasoning here.
Two asks in the meantime: please don't pile on the individuals involved — the process will run its course either way — and if anyone about any signal, report them; they get investigated.
Best
Marc
The signal with 0.0674 is also statistically quite unusual so I recommend investigating it about data reverse engineering and multi accounts( like IP addresses, code characteristics, etc.)
I didn't see 0.0674, but I saw one even higher at 0.08. we will check. but it's also why we have 3 months of live simulation after the end of the contest before declaring a winner. someone who deobfuscates is essentially overfitting, and purely overffit models tend to have a hockey stick performance on live data. for example, rolling sharpe consistently high, then in live period it mysteriously immediately drops. it's happened before and we disqualified the person. other hint in walkforward is test sharpe better than train/validation....which is highly unusual.
It's super hard making things bullet proof, but we don't want to go Numerai approach to force people to buy a token and stake it and get +- token numeraired (no pun intended) tokens based on performance. that's very clever of Numerai but any competent scientist will observe that token performance doesn't match trading performance. we spent months looking at how to solve that decorrelation problem between a utility token performance and trading performance and concluded that to do the tokenomics properly would require heavy regulatory approval.. going from Utility Token to tokenized profitability of trading system is super heavy. maybe it's worth it in the long run. don't know yet.
Alexey and I and some other entreprenuers discussed at length creating an "algorithmic token", which , like algo stablecoins, adjusts via minting and burning a token that has a target that matches trading
Hey TaroKogawa, we did an audit of the 0.08 sharpe signal. It's real. no dataleakage, doesn't appear from validation vs test that contestant completely de-obfuscated the dataset. we'll obviously monitor the signal in the 3 month live period and if it goes hockey stick down that'll be a big red flag, notwithstaning our assessment that it's real. havent checked for multiple accounts yet though.
will revert on the multi accounts and decisions taken from that.
Thanks for being vigilant on this. Why would they choose such a nickname, I wonder!
well, yeah, such a username is a bit interesting...
Thanks to the community member who raised this — the report was taken seriously and a full audit followed.
Here's what we found and what we're doing:
What happened. One participant operated two accounts and submitted duplicate strategies through both, in violation of Rule §6.A (one account per participant). We verified this at the file level and at the prediction level.
What we want to be clear about. The top-ranked signal at the center of the original post was created before the second account existed, without breaking any rules — it was genuine, well-executed work, and the duplication gave it no scoring advantage. The quality-selection process already prevents duplicate signals from double-counting, so no other participant's score or ranking was distorted by the copies.
Why disqualification. The one-account rule exists to keep submission limits meaningful and the leaderboard honest, and it carries disqualification as its stated consequence. The participant was given several opportunities to respond during the review, and we considered their explanation carefully. Weighing everything together, we've disqualified both accounts from the current competition. Their signals will be removed from scoring and the leaderboard soon, and one of the two accounts will be deleted. Any signals that were previously displaced by them will return automatically as the rankings recompute.
What changes going forward. We plan to add automated detection for duplicate submissions across accounts and to review account-linkage signals at registration, so cases like this get caught at submission time rather than by the community after the fact.
We know competitions live or die on trust in the leaderboard. Thank you for holding us to that standard.
Hi everyone. Yes, it was a duplicate account, but I didn't create it to break the rules, I set it up because I had lost access to my email.
Nevertheless, my strategy yielded excellent results, it was developed over the course of three sleepless nights.
As for the vigilant members, I suggest focusing on creating good strategies rather than looking for flaws in others.
I'm glad you like my nickname!
Sign in to reply.