Autoresearch field report · August 9, 2026

Two Kaggle silver medals. Twenty-six days.

A solo, agent-orchestrated campaign reached #34 in NeuroGolf 2026, then crossed into subsurface geology to finish #47 public and #219 private in ROGII Wellbore Geology Prediction. Different domains. The same research system.

2Silver medals
26Campaign days
2Distinct domains
1Solo operator

The evidence

One result can be luck. Cross-domain repetition is a system test.

These competitions reward different kinds of intelligence: one compresses exact visual reasoning into tiny executable graphs; the other reconstructs unseen geology from noisy physical signals. Both were solo entries, completed on short clocks, and both earned silver.

The 2026 NeuroGolf Championship

#34Current final rank

Exact neural program synthesis

The smallest correct ONNX program for each of 400 ARC-AGI tasks, under hidden correctness tests and a hard runtime envelope.

  • Silver medal
  • 9-day sprint
  • 7655.34 public = private
  • 395 submissions
Open leaderboard →

ROGII — Wellbore Geology Prediction

#47Public · #219 private

Physics-grounded geosteering

Predict stratigraphic position along horizontal wells from trajectory, gamma-ray logs, typewell references, and a short interpreted prefix.

  • Silver medal
  • 17-day sprint
  • 5.816 public RMSE
  • 8.256 final private RMSE
Open public leaderboard →

Rank note. Results verified against Kaggle on August 9, 2026. The earlier NeuroGolf post captured the deadline standing at #38; Kaggle's current leaderboard and published write-up title show #34. The ROGII LinkedIn update captured #53 on day 16; the final public standing improved to #47, with #219 on the private leaderboard.

The reusable architecture

Models propose. Executable evidence decides.

Vexorium is not a wrapper around one frontier model. It is a research control plane that turns heterogeneous intelligence into bounded experiments, durable evidence, and safe cumulative progress.

  1. 01

    Formalize

    Freeze the objective, constraints, evaluator, and risk budget.

  2. 02

    Propose

    Route bounded hypotheses across diverse agent and model lanes.

  3. 03

    Execute

    Require runnable artifacts—not persuasive research prose.

  4. 04

    Evaluate

    Test correctness, value, generalization, runtime, and failure modes.

  5. 05

    Remember

    Bank verified wins and preserve refutations, provenance, and floors.

  6. 06

    Transfer

    Challenge the full portfolio whenever one new mechanism works.

  7. 07

    Deploy

    Promote only after external validation; hedge or roll back uncertainty.

Operating principle

Human-directed, evaluator-governed autoresearch: agents supply scale and diversity; explicit objectives, executable tests, and accountable decisions supply direction.

Case study 01 · NeuroGolf

Autonomous grinding found the gains. Portfolio transfer found the missing ideas.

NeuroGolf made every idea executable. A task earned nothing if its ONNX graph failed a hidden case; among correct graphs, smaller parameters and intermediate memory scored higher. That gave the research system an unusually crisp contract: generate, run, measure, and reject.

A dispatcher kept multiple coding-agent lanes working across the 400-task portfolio. Candidates passed local correctness, fresh-generator, crash, and runtime gates before entering a monotonic bank. Independent improvements were pooled into Kaggle submissions, and the protected baseline changed only after the external score improved. Failed ideas became reusable evidence instead of repeated mistakes.

The important distinction was simple: agents proposed; executable evaluators decided.

The official nine-day result moved a public starting artifact from 7240.26 to 7655.34—a gain of 415.08—with identical public and private scores. The starting artifact was not created from zero by Vexorium, so the claim is the measured campaign gain around that baseline, not ownership of every initial point.

The system ran unattended for long stretches, but it was not mythology-level autonomy. Human direction still set strategy, repaired quota and infrastructure failures, and challenged assumptions when the search became repetitive. That boundary is a feature: autonomy should increase experimental throughput without erasing accountability.

Separate research after the deadline

Post-competition work changed the unit of search from task × attempt to mechanism × entire portfolio. A signed-pooling discovery transferred from one task to four more; a finite-precision mechanism added another perfect task. The resulting runtime-safe artifact scored 7672.16 public and private with nine perfect tasks. This is an R&D result, not part of the official medal score.

Read the full Kaggle technical write-up →

Case study 02 · ROGII

The second silver came from knowing when the evaluator could not be trusted.

As the campaign's first LinkedIn field note put it, drilling a horizontal well is like “navigating underground without a map.” The model had to infer true vertical thickness along an unseen lateral from the well path, noisy gamma-ray measurements, a vertical reference log, and a short human-interpreted prefix. Lower RMSE was better.

The operator entered without professional geology training. The system responded by formalizing the physics: gamma ray as a noisy alignment signal, stratigraphic position as an inverse problem, neighboring wells as structural priors, and the interpreted prefix as an anchor. The NeuroGolf loop was adapted rather than copied—fast local cross-validation replaced plentiful leaderboard feedback, while external submissions were treated as scarce experiments.

During the first 9.4 hours, the evaluator-driven loop moved honest cross-validation from 15.91 to 8.399 across 67 bank events and 17 frontier steps. Assumption-breaking lanes ran alongside the grinding loop with predeclared falsification thresholds, closing weak directions before they consumed the remaining days.

The public/private trap

The campaign's best public submission, P95, reached 5.816 and public rank #47—but scored 8.717 on the private wells. A deliberately different second selection, P44, looked much worse publicly at 7.450 yet delivered 8.256 private and the medal. The hedge was not decorative risk language; it was the winning deployment decision.

The retrospective found that paired local/public mechanism directions agreed only 8 times in 22 tests. It also estimated a 0.57-foot RMSE standard deviation for a fresh 150-well draw, while many late improvements were only 0.01–0.08 feet. The system was measuring tiny effects precisely inside one sample, then asking them to survive a much larger shift in wells. More agent throughput could amplify that error; it could not make the target valid.

Superhuman research needs more than scale. It needs objective validity, uncertainty sizing, diverse deployment arms, and the discipline to preserve a result that looks worse on the visible scoreboard.

That lesson strengthens the platform. The first campaign proved persistent autonomous execution. The second proved that orchestration must include epistemics: when to distrust a proxy, when to keep an independent arm, and when final selection matters more than one more leaderboard gain.

The Vexorium moat

The intelligence provider is replaceable. The research memory compounds.

A prior Vexorium post asked whether a team is “cooked” when it cannot access one particular frontier model. These campaigns provide the practical answer: no. The durable advantage is the system around the models.

01

Evaluator engineering

Turn an ambiguous technical ambition into tests that can authorize or reject progress.

02

Provider-agnostic orchestration

Route work across frontier APIs, coding agents, private models, or self-hosted intelligence.

03

Monotonic evidence memory

Preserve wins, failures, counterexamples, provenance, and the exact state that produced them.

04

Mechanism transfer

Convert one breakthrough into a portfolio-wide challenge instead of a one-off anecdote.

05

Autonomous persistence

Keep bounded experiments moving across nights, disconnections, failures, and changing quotas.

06

Governed deployment

Separate discovery from promotion, measure uncertainty, preserve diversity, and roll back safely.

For Vexorium, superhuman intelligence is operational: more useful hypotheses, tested more continuously, remembered more exactly, and transferred more broadly than a person or an ungoverned model could sustain alone.

Put intelligence to work

Do you have a hard problem whose hypotheses can be executed and scored?

Vexorium builds high-intelligence research programs for technical domains where persistent search, exact evidence, and safe deployment matter more than another chat interface.