18 competitions · Unified hidden-test evaluation · Human & LLM comparison
Same template across all 18 competitions; only the five slot contents vary per task.
All 18 GNN-CB competitions across 7 graph categories, 3 difficulty tiers, 4 task types, and 10 domains. Filter by any dimension.
Composition of the 18 GNN-CB competitions attempted by human participants, summarized across four dimensions: graph category, difficulty tier, task type, and application domain. Per-slice counts and competition codes are listed in the legend below each donut chart. See the figure below for additional details.
Human scores vary substantially across competitions, reflecting task difficulty and domain complexity. Easy tasks (C03, C08) achieve high scores with low variance; Hard tasks (C02, C04, C18) show lower averages and larger best–worst gaps.
| # | Competition | Difficulty | N | Metric | AVG ↓ | STD | High | Low |
|---|
We evaluate nine leading LLMs on all 18 GNN-CB competitions using a fixed plan-then-code prompting strategy with a bounded execute-and-repair loop. Across 162 total model–task evaluations, LLMs surpass the Human Top score in only 19 cases, meaning that in more than 88% of evaluations no model reaches the performance achieved by the best human competitor. Overall, the results demonstrate that current LLMs consistently fall short of expert-level GNN competition performance regardless of task difficulty Fig. Overall LLM gap analysis on GNN-CB .
Open-source models generally exhibit larger average gaps and higher variance across competitions, indicating lower consistency and reliability. Proprietary reasoning and chat models perform comparatively similarly, with no clear systematic advantage for reasoning-oriented architectures. Among open-source systems, Kimi-k2.6 approaches the stability of proprietary chat models, while DeepSeek-v4 Pro remains competitive but less consistent on medium and hard tasks.
Easy classification tasks reveal that top proprietary models operate close to human baselines, whereas open-source chat models display larger negative gaps and wider best-to-worst performance spreads. As difficulty increases, mean and worst-case performance degrade substantially across nearly all models, highlighting the limitations of current foundation models under high-complexity graph learning settings. For more analysis see the figures below .
Per-model performance across all 18 competitions under frozen plan-then-code prompting.
| # | Model | Type | Mode | C03 Easy |
C08 Easy |
C01 Med |
C05 Med |
C07 Med |
C09 Med |
C10 Med |
C11 Med |
C12 Med |
C15 Med |
C16 Med |
C02 Hard |
C04 Hard |
C06 Hard |
C13 Hard |
C18 Hard |
|---|
GNN-CB is supported by a globally distributed cohort of contributors from 18 universities across 12 countries and five continents. The benchmark was intentionally designed to reflect broad geographic and institutional diversity spanning Africa, the Americas, Asia, and Europe. This diversity helps ensure that the curated competitions, datasets, and evaluation settings are not centered around a single regional research community, while also providing a heterogeneous human expert baseline.
If you use GNN-CB in your research, please cite: