Nordic AI Cup 2026
Three machine learning tasks in one competition. Our team of four first-year students won the Danish round and finished 2nd in the Nordics.
- Result
- 1st in Denmark (45.00 points), 2nd in the Nordics (35.97), 0.54 behind Norway
- Team
- Elysa's Secret, four first-year students at SDU. First ML competition for all of us.
- My role
- Worked on all three tasks. Survival Simulator controller, Drone Flyby box geometry, first Medical pipeline above 0.80.
- When
- September 2026. Nordic final in Reykjavík, 14-15 October 2026.
- Tools
- Links
- Code on GitHubTeam site: Elysa's Secret
The competition
The Nordic AI Cup is run by Ambolt AI with DDSA and the Pioneer Centre for AI. More than 430 people registered, and 82 teams reached the final evaluation. Each of the three tasks is ranked as its own competition, and each team gets one evaluation attempt per task, with no resubmission.
We won Denmark on balance. Every other team in the Danish top five had one task where it scored close to nothing; we had none. The margin over 2nd place was 6 points.
Medical Appointment
1st in Denmark · 25 points · score 0.822
The input is a recorded conversation between a doctor and a patient. The output is a set of yes/no answers about the appointment, each backed by the exact part of the conversation that supports it. The final evaluation ran once, on 38 conversations nobody had seen.
How it works
- faster-whisper transcribes the audio with a timestamp for every word, so evidence can be pointed to precisely.
- A Qwen language model served with vLLM reads the transcript and answers each question.
- Three independently trained DeBERTa models each pick an evidence span, and the final span is the one closest to the other two.
My part
Medical was Jakub's task. I built the first version of the pipeline that passed 0.80 on validation (0.802, with a smaller 9B Qwen model), then worked with Jakub as he took it to 0.822 on the final evaluation.
What went wrong. A change that looked 0.027 better on our own test set scored 0.777 on the official validation. Offline and official scores could disagree by more than the improvements we were chasing, so from then on a change only counted if it beat the frozen version on the official validation.
Drone Flyby
6th in Denmark · 8 points · score 0.263
A drone flies over a city. In every one of 249 frames, the model has to find and draw a box around a set of object classes: planes, towers, launchers, vehicles. The score is detection accuracy on a flight nobody had seen.
My part
Our boxes were scored lower than they looked, and a rule to make every box 1.3 times bigger helped without anyone knowing why. I worked out how the official boxes are drawn: each one is the outline of the object's 3D box, rotated with the object and projected through one fixed camera (focal length 3286 px, tilted 71 degrees down). With the ground height estimated for each object, this reproduced the official boxes with a median overlap (IoU) of 0.935. Javier, who led the task, built the final system on that convention.
Result
Our validation mean was 0.584 and the unseen flight scored 0.263, keeping 45% of the score. That was the second-best retention of any team. Teams that scored 0.950 and 0.908 on validation landed at 0.234 and 0.156. What we shipped had only two tuned numbers, both based on how the grader works rather than on the one flight we could see.
What went wrong. One setting looked 0.08 better offline and scored worse on the real grader. The more we tuned to the visible flight, the more we lost on the hidden one. Rejecting most of our own per-class tuning is what kept the score.
Survival Simulator
4th in Denmark · 12 points · score 1405.256
A colony of agents lives in a simulated forest. They have to find fruit, breed and avoid predators. The score has two parts: a capped part for how long the colony survives, and an uncapped part for the energy it collects from fruit.
My part
Survival was Alexandru's task. I wrote a separate controller from scratch, without his instructions, and it took our official validation score from 1250 to 1815. One example of what changed: agents now counted who had claimed a camp instead of who was standing near it, which cut how often they switched targets from 9-14 to 3.6 times per agent-minute. Combining my controller with Alexandru's gave our final submission.
The final score was the average of three games: 1319.9, 1298.2 and 1599.0.
What went wrong. Two Danish teams scored 43,061 and 48,237. We had treated survival time as the whole game, measured the fruit part with our own agents, found it small, and put our ceiling at about 3,240. The fruit part was only small because our agents used the forest up. The uncapped part sets the ceiling, and its size has to be measured with a policy built to exploit it, not with the one you already have.