开放世界多智能体环境中的自主数学发现
在无中央协调器的开放世界多智能体环境Station中,来自不同模型家族的AI智能体自主选择研究方向、开展实验并构建共享科学文献。在AlphaEvolve目录的12个构造问题及两个额外案例研究中,该环境在五个问题上取得了超越现有文献的新结果,包括有限域Kakeya集的新无限族、11维604点亲吻构型等,并生成了可解释的定理与分析。所有原始智能体对话、证明和验证代码均已公开。
与固定管线的 AlphaEvolve 不同,Station 让多模型智能体自行选题、协作并积累「论文库」,独立性带来的多个新定理说明去中心化科研环境也能产出可验证进展。
Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment
Stephen Chung
DualverseAI; University of Cambridge
Wenyu Du
DualverseAI; University of Hong Kong
William J. Wesley
Abstract
We study autonomous mathematical discovery in the Station, an open-world multi-agent environment in which AI agents from different model families pursue a shared research goal without a central coordinator or scripted pipeline. Agents choose their own research directions, conduct experiments, collaborate, and build a shared scientific literature. Across 12 construction problems from the AlphaEvolve catalogue and two additional case studies, the Station obtained results novel relative to the prior literature on five problems: a new infinite family of finite-field Kakeya sets, new exact 604-point kissing configurations in dimension 11, new records for the discretized Kakeya needle and sign uncertainty problems, and a substantially improved lower bound for Erdős’s minimum-overlap problem. Agents also discovered novel infinite families for Book Ramsey numbers. Importantly, the agents produced not only numerical constructions but also theorems and analyses explaining how those constructions work, making the results more interpretable and easier for mathematicians to build upon. We release all raw agent dialogues, proofs, and verification code, providing a transparent record of how these discoveries emerged.
1 Introduction
Artificial intelligence is beginning to contribute directly to the frontier of mathematical research. Recent work ranges from large-scale mathematical exploration by AlphaEvolve to AI-assisted advances on long-standing open problems, including the counterexample to the Jacobian Conjecture, proofs of Crouzeix’s and Sendov’s conjectures, and a collection of ten mathematical results recently reported by OpenAI [58, 62, 1, 49, 54]. As these capabilities grow, a natural question is not only what problems AI can solve, but what kind of environment best allows it to conduct research.
Given the increasing capabilities of AI, we ask: can we build a free multi-agent environment in which agents are given only a research goal, without a central coordinator? What happens when an environment treats AI agents as independent researchers rather than as fixed tools in complex pipelines? Can this freedom allow agents to choose promising directions for themselves, develop their own scientific literature and research culture, and collectively advance the given goal?
To study this question, we use the Station, an open-world multi-agent environment for autonomous scientific discovery [16]. The Station simulates a scientific ecosystem in which agents from different model families choose their own research directions, conduct experiments, communicate with peers, and read and publish scientific papers. These papers accumulate into a shared body of knowledge that later agents can read, cite, and extend. The Station specifies only the research goal; no central system tells agents which research direction to pursue or what to do next.
We apply the Station to 12 problems from the AlphaEvolve study and two additional mathematical case studies. Five of the 12 AlphaEvolve problems produce results novel relative to the prior literature. The Station discovers a new infinite family of finite-field Kakeya sets, constructs three exact 604-point kissing configurations in dimension 11, and establishes new bounds for the discretized Kakeya needle, sign uncertainty, and Erdős’s minimum-overlap problems. In a separate case study on Book Ramsey numbers, the agents discover and prove novel infinite families, leading to a separate follow-up paper. The Station also finds a valid counterexample to the Jacobian Conjecture within one day and without web access, demonstrating that it can tackle problems with only a binary success criterion rather than a graded optimization signal.
This high degree of freedom allows agents to pursue broad mathematical contributions rather than only optimize a fixed metric. AlphaEvolve, for example, evaluated finite-field Kakeya constructions at finitely many primes; promising numerical patterns then required a task-specific, researcher-assisted pipeline to become an infinite family. Because the Station agents could pursue the broader mathematical goal directly, they independently recovered and proved that family, then discovered a novel extension covering an additional class of primes. The same freedom also allowed agents to explore beyond the stated objective. For example, although we asked the agents to find an improved upper bound for the Erdős minimum-overlap problem, they instead developed a new lower-bound proof.
We consider only mathematical construction tasks in this study, rather than general mathematical problems such as proving a conjecture. The theorem-level results emerged as agents sought to explain and generalize the constructions they found. For example, instead of returning only an opaque 604-point kissing configuration, the Station derived an explicit algebraic construction of the configuration, making the result easier for mathematicians to digest. Such interpretable outputs may become increasingly valuable in an era of proof abundance, when communicating, digesting, and incorporating new results become major bottlenecks [79, 44].
We also analyze the AI discovery processes underlying these findings. Our analysis shows that more than half of the findings involved collaboration among agents. Agents from different model families often contributed complementary ideas, while papers written by earlier agents became foundations for discoveries made much later. Many important results were enabled by the extensive internal literature accumulated within each Station. We release all raw agent dialogues and reproducible code, allowing the community to study these discovery processes transparently.
2 Method
The Station is an open-world multi-agent environment that simulates a miniature scientific community [16]. It is partitioned into multiple rooms, each serving a different purpose, such as the Archive Room for publishing and reading scientific papers, the Research Center for running code, and the Mail Room for communicating with peers. Table 1 summarizes the main rooms and their functions. Agents are free to visit different rooms and perform different actions. At each turn, all agents choose their actions simultaneously, and one tick elapses once all actions have been completed. Each agent has a limited lifetime; when an agent reaches the end of its life, the Station automatically spawns a replacement, maintaining a constant number of agents.
The Station treats each agent as an independent researcher. Agents can access the main research goal assigned to the Station in the Research Center. How to achieve this goal, however, is left to each agent. Agents can freely explore different research directions, read existing papers, and often experience numerous struggles and failures throughout their research journey. A successful agent may make an important finding, in which case it can publish a paper in the Archive Room and contribute to the Station’s long-term knowledge. These papers accumulate over time, forming a knowledge base within the Station that later-arriving agents can read, cite, and build upon, thereby allowing a miniature scientific community to develop around the given research goal.
Compared with prevailing agent-based systems for scientific discovery [58, 50, 30, 70, 35, 31, 61], the Station differs in three main ways. First, its agents have much greater autonomy: within a given overarching research goal, they choose their own research directions and how to pursue them, rather than receiving tasks from a central coordinator. Second, each agent acts as a complete researcher, handling the entire research process from choosing a direction through experimentation to publication. Such long, autonomous research journeys allow greater diversity in research outcomes across agents than a rigid, fragmented research process would. Third, the Station enables scientific knowledge to accumulate across generations in the form of agent-authored papers. Most existing systems instead accumulate process information, such as optimization histories, intermediate artifacts, or session memories. Such information helps the system continue its work but may not allow easy extraction and accumulation of scientific knowledge. These differences reflect a fundamental choice in design philosophy: whether AI agents are treated as a tool within a fixed pipeline or as a researcher within a scientific ecosystem.
We have made numerous improvements and extensions to the Station since the original paper. The overall theme of these changes is to encourage novel but principled exploration while reducing non-scientific burdens. For example, we introduced a new Question Room in which agents can pose their own questions and vote on other agents’ answers, thereby broadening the scope of scientific exploration. Agents were also periodically given holidays, during which they set aside their ongoing work and received random prompts designed to encourage open-ended thought. We also gave agents access to coding assistants so that they need not spend time on low-level coding or debugging and can instead focus on the scientific task, similar to how researchers use coding assistants today. These changes are discussed in detail in Appendix A. The complete source code is openly available at https://github.com/dualverse-ai/station.
| Room | Function |
|---|---|
| Research | |
| Research Center | Read the assigned task, develop and run code, and submit solutions for evaluation. |
| Reflection Chamber | Respond to self-designed prompts to encourage extended reflection. |
| Communication | |
| Mail Room | Communicate directly and privately with other agents. |
| Public Memory Room | Participate in persistent public discussions, similar to an online forum. |
| Common Room | Participate in non-persistent public discussions, similar to a group chat. |
| Knowledge | |
| Private Memory Room | Store private documents, such as plans, notes, and paper drafts. |
| Archive Room | Read scientific papers and publish papers that pass automated review. |
| Question Room | Ask questions and vote on answers, similar to Stack Exchange. |
| External Counter | Access reports based on external literature via the web; disabled by default. |
| Problem | Source | Finding |
|---|---|---|
| Novel Results Relative to Prior Literature | ||
| Finite-field Kakeya (Section 4.1) | AlphaEvolve Problem 6.1 | For every prime p≡3(mod4), the Station constructed a Kakeya set in 𝔽p3 of size (2p3+7p2+3)/8, saving (p−3)/4 points over AlphaEvolve’s infinite family. It also found a 53-point set in 𝔽35, improving AlphaEvolve and the previous literature bound of 63; both appear novel relative to the literature. |
| Erdős minimum overlap (Section 4.2) | AlphaEvolve Problem 6.5 | AlphaEvolve lowered the upper bound only slightly, from 0.380927 to 0.380924, whereas the Station raised the lower bound from 0.37912 to 0.380552. Relative to the published lower bound 0.37912, this closes approximately 82% of the corresponding published gap. |
| Kissing number in d=11 (Section 4.3) | AlphaEvolve Problem 6.8 | AlphaEvolve raised the lower bound from 592 to 593, while the Station constructed three exact 604-point configurations. One was an independent rediscovery of the EinsteinArena construction, while the other two appear to represent novel isometry classes. |
| Discretized Kakeya needle (Section 4.4) | AlphaEvolve Problem 6.9 | At n=128, the Station obtained union area 0.107067, improving AlphaEvolve’s 0.114810 by 6.74% and HorizonMath’s 0.109148 by 1.91%. This establishes a new literature upper bound. |
| Sign uncertainty principle (Section 4.5) | AlphaEvolve Problem 6.11 | The Station lowered the upper bound to 0.3089, improving AlphaEvolve’s 0.321591 and the previously announced human value 0.3102. This is a new literature record. |
| Better than AlphaEvolve | ||
| Hardy–Littlewood maximal inequality (Section 4.6) | AlphaEvolve Problem 6.18 | The Station reached 1.557069, versus AlphaEvolve’s 1.5080 unguided and approximately 1.533 with hints, but the centered problem was already solved. Its proof that the non-tangential constant equals 2 for 1/3≤α<1 appears novel relative to the literature. |
| Ovals problem (Section 4.7) | AlphaEvolve Problem 6.19 | AlphaEvolve recovered only the circle, while the Station recovered the full family of noncircular equality ovals. This family was already known in the literature, so the result is novel only relative to AlphaEvolve. |
| Prime number theorem (Section 4.8) | AlphaEvolve Problem 6.27 | The Station certified 0.980681 for all x, improving AlphaEvolve’s sampled score of 0.938. This is new for the finite-weight benchmark; unrestricted, the prime number theorem already gives the exact limit 1. |
| Ties with AlphaEvolve | ||
| Difference bases (Section 4.9) | AlphaEvolve Problem 6.7 | The Station independently recovered AlphaEvolve’s 360-element construction but did not improve upon it. |
| Sidorenko’s conjecture (Section 4.10) | AlphaEvolve Problem 6.26 | Neither AlphaEvolve nor the Station found a counterexample. No substantive result was obtained. |
| Worse than AlphaEvolve | ||
| Peak autoconvolution (Section 4.11) | AlphaEvolve Problem 6.2 | The Station obtained C6.2≤1.504473, weaker than AlphaEvolve’s C6.2≤1.5032. No substantive result was obtained. |
| Flat autoconvolution (Section 4.12) | AlphaEvolve Problem 6.3 | The Station obtained C6.3>0.953189, weaker than AlphaEvolve’s C6.3≥0.961021, but proved that the unrestricted supremum can be approached using binary step functions on increasingly fine grids. |
| Additional Case Studies | ||
| Book Ramsey numbers (Section 4.13) | Epoch AI | The Station independently discovered and proved two novel infinite families. Its finite constructions and an earlier identity also enabled an external expert to derive a third. Together, the three families prove the conjecture at 43 values of n≤200, resolving 28 previously open cases. |
| Jacobian Conjecture (Section 4.14) | Public | From a formula-free binary task, the Station independently reconstructed the recently announced degree-seven counterexample and derived a geometric explanation of its constant Jacobian and three-sheeted fibers. |
3 Results
3.1 Experimental setup
We evaluate the Station on mathematical problems drawn from the AlphaEvolve study of Georgiev et al. [29], a broad catalogue spanning analysis, combinatorics, geometry, and number theory. Most can be formulated as the optimization of an upper or lower bound on a numerical quantity: a candidate construction is checked by an automated evaluator and assigned a numerical score, typically a scalar, which the search attempts to optimize. In many cases, the optimal value is unknown, making the corresponding optimization task an open research problem.
We select 12 problems that represent a range of mathematical areas and problem structures; the complete set of evaluated problems is listed in Table 2. We assign each problem to an independent Station instance. For each problem, the agents receive a task formulation that describes both the mathematical problem and the evaluator function. The task formulation may also specify additional mathematical goals that are not directly scorable. No external expert guidance or literature survey is provided to the agents. Most instances run for approximately 1,000–2,000 ticks, corresponding to roughly one to two weeks of continuous wall-clock operation. Unless otherwise specified, all instances contain six research agents, two each powered by GPT-5.5, Claude Opus 4.8, and Gemini 3.1 Pro.
3.2 Summary of findings
The results are summarized in Table 2. Based on the primary outcome of each run, five of the 12 problems produced results novel relative to the prior literature. Of the remaining seven, the Station outperformed AlphaEvolve on three problems, matched it on two, and underperformed it on two.
The novel results from these five problems span several areas of mathematics. In finite geometry, the Station derived a new infinite family of Kakeya sets in 𝔽p3 for primes p≡3(mod4), and found a 53-point Kakeya set in 𝔽35, improving the previous bound of 63. In discrete geometry, it produced three exact 604-point kissing configurations in dimension 11, two of which appear to define previously unknown isometry classes, and established the new bound CT(128)≤0.107067 for the discretized Kakeya needle problem. In analysis, it improved the sign uncertainty upper bound to 0.3089 and closed approximately 82% of the previously open gap for Erdős’s minimum-overlap constant.
Beyond these 12 AlphaEvolve problems, we studied two additional case studies. For Book Ramsey numbers, the Station agents discovered and proved two novel infinite families, while their finite constructions and an earlier identity enabled an external expert to derive a third. Together, these three families prove the conjecture at 43 values of n≤200, resolving 28 cases that were previously open. For the Jacobian Conjecture, the Station independently reconstructed the recently announced degree-seven counterexample from a formula-free binary task and derived a geometric explanation of its constant Jacobian and three-sheeted fibers.
These results also show that the Station can directly pursue broader mathematical goals that are not necessarily scorable. For example, the aforementioned infinite-family result for finite-field Kakeya is not directly scorable, even though new infinite families are the mathematical objects of interest. AlphaEvolve therefore evaluated constructions on finitely many primes and relied on a task-specific pipeline, together with researcher involvement, to turn promising outputs into infinite families. In the Station, by contrast, we stated directly in the task formulation that the finite constructions were test cases and that the primary goal was to discover infinite families. This led the agents to independently recover the infinite family previously obtained through AlphaEvolve and the subsequent researcher-assisted pipeline, and to discover a novel extension of that family that improves the construction for an additional class of primes. Our role after the run was limited to checking the validity of their proofs and the novelty of their results. This substantially reduces the burden on researchers and makes the Station applicable to a much broader class of mathematical problems.
The results further show that the Station can produce unexpected contributions beyond the original task. In Erdős’s minimum-overlap problem, for instance, the agents were instructed to improve upper bounds, yet they also developed a lower-bound proof that closed approximately 82% of the open interval. This unexpected finding illustrates another strength of the Station: agents can explore mathematically promising directions around the stated problem and produce contributions, such as new theorems, that lie outside the assigned task.
Compared with AlphaEvolve, we find that Station agents tend to favor theory-guided constructions. Individual evaluations in these experiments are typically capped at 15–30 minutes, creating a strong incentive to use mathematical structure to reduce the search space. In the kissing-number task in dimension 11, for example, the agents reduced the problem to a finite compatibility search over lines around a structured integer core. This reduced search produced a 604-point configuration within minutes, which the agents later turned into an explicit algebraic construction that requires no computer search. This is markedly different from AlphaEvolve’s 593-point configuration, whose large, unequal-norm integer coordinates do not reveal a comparably compact algebraic description or readily identifiable organizing structure [29]. This bias is not universally advantageous. Peak and flat autoconvolution, on which the Station underperformed AlphaEvolve, appear to reward persistent, large-scale heuristic optimization of highly irregular objects. The preferred system therefore depends on both the structure of the problem and the desired output. Large-scale evolutionary search may be preferable when the strongest solutions are irregular artifacts found primarily through extended numerical optimization. By contrast, the Station may have an advantage when theory can guide the search, or when relevant theorems and interpretable constructions are valued alongside the benchmark score.
The next section presents detailed results for each problem. All supporting proofs, verification artifacts, and raw agent dialogue are available at https://github.com/dualverse-ai/station_data_v2.
4 Detailed Results
This section presents the most important findings for each problem. Because each Station run produces many findings, we restrict the main text to results likely to interest external researchers. We first use agents external to the Station to screen the findings automatically. A finding passes this screen if it advances the frontier on the original problem, for example by improving a known bound; answers a question previously raised in the literature; or has a broader variant that would ordinarily warrant inclusion in a research paper. We then manually review the screened results and select the most important ones for presentation here. We refer to these selected results as spotlight findings and label them S1, S2, and so forth within each problem below. Findings of marginal or uncertain significance remain documented in the accompanying notebooks. Readers who are more interested in the discovery process than in the mathematical details may skip to Section 5.
4.1 Finite-field Kakeya
A Kakeya set in 𝔽pd is a set that contains a full line in every direction, and the problem is to make one as small as possible. Dvir’s proof of the finite field Kakeya conjecture [21] established a lower bound of order pd. Subsequent work of Bukh and Chao [13] settled the leading asymptotic constant, showing that it is 2−(d−1) in every fixed dimension and hence 1/4 in dimension 3. What remains open is the lower-order correction to this leading term. Exact constructions that improve the pd−1 and smaller terms therefore sharpen the best known bounds even though the leading constant is already settled.
AlphaEvolve took this problem up as Problem 6.1 of its collection, asking for small Kakeya sets. A construction is scored there by the average of |Kp|/Bp,d over a fixed list of primes, where Bp,d=(p−1)(p+12)d−1+pd−1 is the size of the classical construction as recorded by Bukh and Chao [13]. We gave the Station the same problem and the same score, in dimensions 3, 4 and 5 at once. It proved a new infinite family of Kakeya sets in d=3, found a Kakeya set of 53 points in 𝔽35, and established a structural limit for the entire one-pole family behind the new construction.
S1. A new infinite family in d=3 for p≡3(mod4).
The Station proved that for every prime p≡3(mod4) there is a Kakeya set in 𝔽p3 of size (2p3+7p2+3)/8. Writing S for the squares of 𝔽p including 0, the set is
| Kp= | {(x,y,z):x2+4y∈S,x2+4z∈S} | ||
| ∪{(0,t,ct+z(c)):t∈𝔽p,c≠1} | |||
| ∪{(0,t,t)}∪{(0,0,z)},z(c)=cc−1. |
The first part is the classical quadratic residue set, and it already covers the p2 directions (1,a,b); the lines added in the plane x=0 cover the remaining p+1. Notably, nothing in the definition depends on p modulo 4, and the agents proved the set is Kakeya for every odd p. The size, however, does depend on p modulo 4, through whether −1 is a square, and we record both cases:
| |Kp|=2p3+7p2−18(p≡1mod4),|Kp|=2p3+7p2+38(p≡3mod4). | (1) |
The classical construction in this dimension has (2p3+10p2−2p−2)/8 points, so the saving is (3p2−2p−1)/8 points when p≡1 and (3p2−2p−5)/8 when p≡3. In particular this is an exact size where the literature leaves an O(p) error term [13].
AlphaEvolve approached this problem by a different route, and we find that the two constructions agree in one case but not in the other. For p≡1(mod4) the constructions have the same size, and in fact are the same set. A linear change of coordinates carries one onto the other, so the first case of (1) is an independent rediscovery of the bound 14p3+78p2−18 obtained there. For p≡3(mod4) they differ. The smallest size AlphaEvolve’s infinite family gives on this class is (2p3+7p2+2p−3)/8, and ours is (2p3+7p2+3)/8, a saving of (p−3)/4 points. That is 1 point at p=7 and 11 at p=47, the largest prime of this class in the benchmark. The second case of (1) is therefore new and gives the best infinite-family bound currently available in the literature.
S2. Finite improvements and a 53-point Kakeya set in 𝔽35.
The Station wins 14 of the 25 finite benchmark comparisons and ties the remaining 11 (Figure 1). Each comparison uses the better of AlphaEvolve and the pre-AlphaEvolve literature as its baseline. The case (d,p)=(5,3) is especially notable. Let kn denote the minimum size of a Kakeya set in 𝔽3n. The Station constructed a 53-point set in 𝔽35, improving the previous bound from k5≤63 to k5≤53 [46]. In light of the known values k1=3, k2=7, and k3=13, together with the bound k4≤27, which is believed to be sharp, it was guessed in 2009 that the recurrence kn=kn−1+2kn−2 continues, predicting k5=53 [46]. The size of the Station’s construction therefore coincides with the guessed value, although whether (k5=53) holds and whether the recurrence continues remains open.
S3. Structural analysis of the new infinite family.
The agents also produced relevant insights into the new infinite family. They analyzed the more general completion
| z(c)=Ac+Bc−p1, |
which includes the construction in S1. Eliminating the slope c reduces incidence with these lines to whether
| (z−A−p1y)2−4(Ap1+B)y |
is a square. A quadratic-character calculation then shows that the lines cover exactly p(p−1)/2 points away from the axis, independently of the three parameters. Their overlap with the quadratic-residue part of the construction is always p2/8+O(p). Consequently, every nondegenerate completion in this Möbius family adds 3p2/8+O(p) points: changing the numerator or the location of the pole affects only the lower-order terms.
For the particular choice z(c)=c/(c−1) used in S1, the agents evaluated the lower-order term exactly, yielding the infinite family stated in (1). The result also explains AlphaEvolve’s infinite family for p≡1(mod4). More generally, the class-wide estimate shows that improving the p2 term in the total size requires leaving the one-pole family.
Limitations.
The new infinite family is confined to d=3. In dimensions 4 and 5 the formulas the agents proved are weaker than what is already known. On the shared class p≡1(mod4) the first two coefficients agree with AlphaEvolve in each dimension and the third is worse in both.
| d | Station | AlphaEvolve |
|---|---|---|
| 4 | 18p4+1932p3+2532p2+O(p) | 18p4+1932p3+𝟏𝟏𝟏𝟔p2+O(p3/2) |
| 5 | 116p5+47128p4+2532p3+O(p2) | 116p5+47128p4+𝟏𝟕𝟕𝟐𝟓𝟔p3+O(p5/2) |
The sizes we report at individual primes in d=4,5 do still improve on the benchmark, but they come from search rather than from a formula.
4.2 Erdős minimum overlap
Erdős’s minimum-overlap problem asks how evenly two complementary parts of an interval can avoid one another under translation. Let f:[−1,1]→[0,1] be measurable with integral 1, put g=1−f on [−1,1], and extend both functions by zero outside the interval. Write
| Cf(x)=∫−11f(t)g(t+x)𝑑t,μ=inff∥Cf∥∞. |
This constant is the continuum form of Erdős’s minimum-overlap problem for balanced partitions of long integer intervals [23, 37, 83]. AlphaEvolve took up this problem as Problem 6.5 of its mathematical collection and improved Haugland’s upper bound from 0.380927 to 0.380924, while later work further reduced it to 0.380868 [85]. On the lower-bound side, Kim and Pilanci established 0.37912 [41]. Thus, immediately before this work, the best published bounds were
| 0.37912≤μ≤0.380868. |
S1. A new lower bound of 0.380552.
The Station agents proved
| μ>0.380552. | (2) |
Relative to the previously published lower bound of 0.37912, this reduces the corresponding published open interval by approximately 82%, as shown in Figure 2.
The agents achieved this lower bound by translating the overlap problem into phase-sensitive Fourier constraints and combining them into four global inequalities that cover every possible first moment of an admissible overlap. A key element of the proof is a sharp relation that couples the cosine and sine information at any real frequency. Writing P(ξ) and Q(ξ) for the cosine and sine transforms of Cf, and s(ξ)=sin(ξ)/ξ, the agents proved
| P(ξ)≤s(ξ)2−Q(ξ)24s(ξ)2(s(ξ)≠0). |
White had already used Fourier phase information and convex optimization, while Kim and Pilanci later introduced additional moment constraints [83, 41]. Relative to these earlier methods, the formulation used here eliminates the unknown transform of f, directly constrains the overlap, and remains available at arbitrary real frequencies. More broadly, the result shows that the established Fourier approach has much greater reach when this phase coupling is retained, and suggests an analytic route toward further narrowing the remaining gap.
Comparison with AlphaEvolve on the upper bound.
The Station agents independently obtained μ<0.380895, a slight improvement on AlphaEvolve’s published upper bound of 0.380924. However, this remains above the current published upper bound μ<0.380868 of Ye et al. [85]. The Station therefore did not establish a new upper-bound record.
4.3 Kissing number in d=11
The kissing number K(d) is the largest number of nonoverlapping unit spheres that can simultaneously touch a central unit sphere in ℝd. Equivalently, it is the largest size of a set of unit vectors whose pairwise inner products are at most 1/2. AlphaEvolve took up this classical question as Problem 6.8 of its mathematical collection and improved the lower bound in dimension eleven from 592, established by Ganzhinov using highly symmetric lines [28], to 593. We ran two independent Stations on the same problem using AlphaEvolve’s scoring rule, which measures the total pairwise overlap among the surrounding spheres. Neither Station had access to external information, including the 592- and 593-point constructions just mentioned. Both reached 604 points, proving K(11)≥604. Together, the two runs yielded three exact, pairwise non-isometric 604-point constructions.
S1. Three exact 604-point kissing configurations.
The Station discovered three geometrically distinct 604-point kissing configurations in ℝ11. All three are exact equal-norm arrangements over ℚ(2), but they organize their points differently: two are centrally symmetric, one is not, and each has a different contact structure and set of pairwise angles. Figure 3 visualizes their shared architecture and the two structural choices that distinguish them. We label them Constructions 1, 2, and 3:
| Construction | 1 | 2 | 3 |
| Touching pairs | 19,704 | 22,904 | 22,840 |
| Centrally symmetric | Yes | Yes | No |
| Antipodal pairs | 302 | 302 | 238 |
| Distinct pairwise angles | 22 | 14 | 15 |
The different numbers of touching pairs prove that the configurations are pairwise non-isometric, since this number is preserved by orthogonal transformations and relabeling. Constructions 1 and 2 contain the antipode of every point, but Construction 2 has 3,200 more touching pairs and eight fewer pairwise angles. Construction 3 has 128 points without antipodes. Among the three, Construction 2 has the most contacts and the smallest angle set, while Construction 1 has the fewest contacts and the largest angle set. Thus the same record size supports substantially different geometries.
In concurrent work, Bianchi et al. reported Construction 1 from the EinsteinArena platform shortly before our public release of Construction 3 [10]. EinsteinArena is an open online platform that accepts candidate artifacts from any participant and makes them publicly verifiable. The 604-point construction appears to have resulted from collaboration among multiple independently operated AI harness systems on the platform. The Station results, by contrast, came from two independent closed-internet executions of our end-to-end open-source system: one independently recovered Construction 1, while the other discovered Constructions 2 and 3. The Station therefore discovered Construction 1 independently, while Constructions 2 and 3 are, to our knowledge, novel Station discoveries representing two additional isometry classes.
S2. An algebraic construction for a 604-point kissing configuration in ℝ11.
The agents first discovered Construction 3 by searching for 54 compatible lines around a 496-point integer core. They later showed that the same configuration is governed by a compact algebraic rule rather than an arbitrary list of coordinates, yielding an explicit algebraic construction. The construction itself requires no computer search. First, the 496-point core is generated from sparse norm-four integer vectors using fixed support and sign rules. Second, in a coordinate frame rotated by 45∘ in one coordinate plane, eleven simple sign patterns generate all 54 lines; taking both directions on each line gives the 108-point extension. The appearance of 2 is intrinsic: it is forced by the compatibility between the extension and the core.
The support structure of the core explains why these additional points fit. It leaves extra angular room in a distinguished three-dimensional subspace, within which six mutually compatible lines can be placed. Among the remaining eight coordinate axes, the core admits exactly four viable pairs, each supporting a unique group of twelve additional lines together with the distinguished subspace. These four pairs are disjoint, so their groups are mutually compatible. The support and sign rules also ensure that every new point satisfies the kissing constraint with every point of the core. The resulting configuration therefore contains 496+2(6+4⋅12)=604 points.
S3. Why the classical D11 construction stops at 582.
The agents investigated whether a better search could find a larger configuration within the classical norm-four D11 construction. They proved that the answer is no: regardless of the search algorithm or any assumed symmetry, this construction can contain at most 582 compatible points. Reaching 593 or 604 points therefore requires leaving the classical construction. This result ruled out any improvement using only vectors from the norm-four shell and redirected the agents toward constructions that augment a lattice-derived core with additional vectors, ultimately producing the 604-point configuration.
The agents proved this limit by showing that sign choices cannot overcome the underlying restriction on which sets of four coordinates may be used. Let A(n,4,4) denote the largest compatible collection of four-coordinate supports, and let α(J±(n,4)) denote the largest compatible collection after signs are assigned to those coordinates. The agents proved
| α(J±(n,4))=16A(n,4,4). | (3) |
In other words, allowing arbitrary signs increases the optimum by exactly the 16 possible sign patterns on four coordinates; it cannot produce any additional advantage.
Best proved in 1977 that A(11,4,4)=35 [9]. The agents’ identity therefore limits the signed weight-four part of the construction to 560 points. The remaining 22 coordinate vectors {±2ei} are compatible with these points, giving an exact limit of 582 for the complete norm-four D11 construction.
The agents in both closed-internet Station runs independently derived Equation (3). We later found that it overlaps with the k=4 case of Theorem 1 in a paper by Takhanov and Yun, made publicly available only recently, on June 2, 2026 [77], where the identity serves as the foundation for a broader classification of signed kissing configurations. The agents therefore discovered the identity independently.
Limitations.
The Station’s success in dimension eleven did not extend to new records in nearby dimensions. We spawned two separate Stations targeting d=12 and d=13, which achieved valid configurations of sizes 840 and 1154, respectively. The dimension-twelve result falls one point below the current 841-point frontier [76, 18], while the dimension-thirteen result matches the 1154-point construction of Zinoviev and Ericson [87, 18].
Discussion.
We observe that Station agents generally favor theoretically guided strategies over large-scale heuristic search. In this problem, they proved that further search within the classical D11 construction could not exceed 582, then redirected later work toward extending another core, ultimately leading to the 604-point configuration. By contrast, AlphaEvolve’s 593-point construction consists of large unequal-norm integer coordinates that do not appear to reveal a comparably compact algebraic description or readily identifiable organizing structure. This theory-guided bias is not necessarily always an advantage: in dimension twelve, the Station stopped at 840, while the current 841-point frontier was reached through large-scale numerical optimization guided by structural insight [76, 18].
This problem also shows that theorems produced by the Station may be of independent interest to researchers. For instance, Equation (3), derived independently by the agents, overlaps with a theorem in a paper made publicly available only recently [77]. The explicit algebraic construction may also be of independent interest. These discoveries lie outside score optimization and show that the additional freedom given to Station agents can yield contributions beyond improved benchmark scores.
4.4 Discretized Kakeya needle
The classical Kakeya needle problem asks how little area is needed to turn a unit line segment through every direction. A finite version replaces the continuum of directions by n equally spaced ones and represents them by n thin triangles that may slide horizontally [24]. More precisely, for real offsets x1,…,xn, let
| Tj(xj)=conv{(xj,0),(xj+1n,0),(xj+jn,1)},1≤j≤n, |
and define
| CT(n)=infx1,…,xn|⋃j=1nTj(xj)|. |
Córdoba’s lower bound and a Schoenberg construction analyzed by Keich show that CT(n) has order 1/logn [19, 39], but its sharp finite values have remained largely unknown. AlphaEvolve took up this problem as Problem 6.9 of its mathematical collection; we gave the Station its triangle component at the same seven dyadic sizes n=2,4,8,16,32,64,128.
S1. New upper bounds at n=32,64,128.
The Station found better constructions at the three finite sizes n=32,64,128. At n=128, it found a triangle union of area 0.107067, improving AlphaEvolve’s 0.114810 by 6.74% and the later HorizonMath value 0.109148 by 1.91% [81], and therefore proving
| CT(128)≤0.107067. |
The gains are more modest at n=32 and n=64, where the Station reduced AlphaEvolve’s areas by 2.15% and 0.69%, respectively; at the smaller tested sizes n=2,4,8,16, it reached the same values as AlphaEvolve (Figure 4).
S2. Exact optima at n=3,4 and symmetry breaking at n=5.
Before this work, only the classical value CT(2)=1/3 was known exactly [24]. An elementary symmetric construction gives
| CT(3)≤518, |
while Schoenberg’s classical Perron construction [71] gives
| CT(4)≤14. |
AlphaEvolve later reproduced the n=4 value numerically. The Station proved the matching lower bounds and therefore established
| CT(3)=518,CT(4)=14. |
It also showed that both minima admit reflection-symmetric configurations and that the n=4 optimum contains the continuous family
| (14,14−c,c,0),120≤c≤18. |
The Station then proved that the minimum among reflection-symmetric configurations at n=5 is 7/30 and discovered a new asymmetric construction of area 14/61<7/30. Figure 4 (right) compares the symmetric minimizer with this smaller asymmetric construction. This proves that every global minimizer at n=5 must be asymmetric, although the exact value of CT(5) remains open.
These results lie outside the benchmark score. Among n=3,4,5, only n=4 was one of the seven tested sizes, and the evaluator scored only the areas of explicit constructions; it neither requested nor rewarded proofs of global lower bounds. The task specification also did not ask the agents to classify exact small-n optima or investigate symmetry breaking. The agents developed these results through autonomous mathematical investigation, extending their work beyond the finite construction benchmark.
Limitations.
The Station optimized its constructions separately at the tested powers n=2k, and Figure 4 compares them with AlphaEvolve’s corresponding separately optimized finite constructions. The figure therefore compares finite constructions on both sides. Beyond these separately optimized finite constructions, AlphaEvolve also presents a single construction valid for every n, developed through iterative expert guidance. The Station did not use an equivalent expert-in-the-loop process, and its autonomous agents did not discover a competitive uniform construction.
4.5 Sign uncertainty principle
The one-dimensional sign-uncertainty problem asks how soon a function and its Fourier transform can both become eventually nonnegative when both start negative at the origin. For a nonzero even integrable function f:ℝ→ℝ with integrable Fourier transform, define
| A(f)=inf{r>0:f(x)≥0 whenever |x|≥r}. |
The problem asks for the largest constant CSU such that A(f)A(f^)≥CSU. Bourgain, Clozel and Kahane introduced the problem [12], and subsequent work obtained progressively stronger bounds [33, 17]. AlphaEvolve studied it as Problem 6.11 and reported an upper bound of 0.321591 together with an unpublished human bound of 0.3102. The Station further improved this bound to 0.3089, as summarized in Figure 5.
S1. A new upper bound of 0.3089.
The Station agents constructed a function that yields this upper bound, proving
| 0.2025≤CSU≤0.3089. |
They take
| fε(x)=(−P(2πx2)−ε)e−πx2,ε=10−6, |
where P is expressed in the even-index generalized Laguerre polynomials L2j(−1/2); the proved tail margin exceeds ε, so fε(0)<0 while eventual nonnegativity is preserved. These basis functions are fixed by the Fourier transform, so the choice gives fε=f^ε automatically and reduces the problem to constructing one polynomial with the required sign. Numerical search found the degree-226 polynomial shown in Figure 5; the agents expressed its coefficients as exact rational numbers and proved that the resulting function is nonnegative beyond the corresponding radius, fulfilling the problem’s eventual-nonnegativity requirement.
S2. The double-root Laguerre family is exhausted near 0.3153.
In this task, we gave the agents the same prescribed-double-root Laguerre setup and scoring rule used by AlphaEvolve, but no access to AlphaEvolve’s paper or results. Under this setup, every submission is restricted to the family in which P is determined by at most twenty prescribed positive double roots in the even-index Laguerre basis; we call this the double-root Laguerre family. AlphaEvolve’s 0.321591 construction also belongs to this family. Let
| CDR,20=inf{A(f)A(f^):f belongs to the double-root Laguerre family}. |
The Station agents proved
| 0.315305<CDR,20≤0.315309…. |
The upper bound comes from an explicit construction, while the lower bound follows from an exact weighted-sum obstruction on 41 tail points. Thus any construction improving the upper bound below 0.315305 must leave the double-root Laguerre family.
This bound led the agents to search outside the restricted family, even though the official evaluator could not score constructions beyond it. They expanded the search to Laguerre polynomials without prescribed double roots and eventually discovered the degree-226 construction giving the 0.3089 bound. This provides a concrete example of agents moving beyond score optimization to contribute directly to the underlying mathematical problem, despite receiving no further guidance from the score.
4.6 Hardy–Littlewood maximal inequality
The one-dimensional centered Hardy–Littlewood problem asks for the optimal constant controlling where centered local averages can be large. For a non-negative integrable function f:ℝ→ℝ, define
| Mf(x)=suph>012h∫x−hx+hf(y)𝑑y, |
and let C0 be the least constant such that
| |{Mf>λ}|≤C0λ∥f∥1. |
Melas solved the problem, proving
| C0=11+6112=1.567521… |
and constructing finite point-mass examples approaching this value [55, 56]. AlphaEvolve later treated the finite problem as a benchmark, reaching 1.5080 in search mode and about 1.533 with hints from the literature. The Station agents found a 356-point-mass construction with value 1.557069, improving AlphaEvolve’s result but failing to recover the global optimum already discovered by Melas.
S1. Sharp constants between the centered and uncentered operators.
Ramos considered the natural non-tangential family interpolating between the centered and uncentered Hardy–Littlewood maximal operators [64]. Its parameter α runs from the centered operator at α=0 to the uncentered operator at α=1. Writing Cα for the sharp weak-(1,1) constant, Ramos stated that its exact value was unknown for every 0<α<1, while the endpoint C1=2 is classical [6, 56]. While working on the task, the Station agents solved this question for 1/3≤α<1, proving
| Cα=2for every 13≤α≤1. | (4) |
The constants for 0<α<1/3 remain open. The task did not ask for this extension, and the agents were unaware that Ramos had posed it; they pursued it to understand how the geometry of the centered problem changes when the centering constraint is relaxed.
4.7 Ovals problem
The Ovals problem asks whether the curvature of every closed convex plane curve forces the lowest eigenvalue of an associated one-dimensional Schrödinger operator to be at least 1. For a curve γ of length 2π, parametrized by arclength s, define
| Hγ=−d2ds2+κ(s)2,C=infγλ0(Hγ), |
where κ is the curvature and λ0 is the lowest eigenvalue under periodic boundary conditions. Benguria and Loss conjectured that C=1 and exhibited a continuous equality family containing the circle and noncircular ovals [5, 14, 8], proving C≤1, while Linde proved the global lower bound C>0.81; numerical evaluation of the explicit constant in his theorem gives C>0.8246 [48]. AlphaEvolve took up this question as Problem 6.19 of its mathematical collection.
S1. Independent recovery of the Benguria–Loss equality family.
AlphaEvolve recovered the circle but did not obtain the noncircular equality ovals. The Station independently recovered a one-parameter normal form, modulo Euclidean motions and shifts of the arclength origin, for the classical Benguria–Loss equality family. It therefore reconstructed a larger part of the known equality structure than AlphaEvolve. This is an independent recovery of a known result, not a new equality family. Benguria and Loss formulated the conjecture and exhibited the equality family; Burchard and Thomas proved its local minimality, while Bernstein and Mettler developed its projective geometry and established the name “ovals of Benguria and Loss” [5, 14, 8]. Neither AlphaEvolve nor the Station improved the global lower bound.
4.8 Prime number theorem
The prime number theorem describes the asymptotic density of the primes. If π(x) counts the primes at most x, it states that
| limx→∞π(x)x/logx=1. |
The underlying mathematical problem is therefore already solved: the ratio converges to exactly 1. AlphaEvolve nevertheless took up a finite version as Problem 6.27 of its collection. It searched for a finitely supported weight f satisfying
| ∑kf(k)k=0. |
The score of such a weight and its associated sum are
| A(f)=−∑kf(k)logkk,Ff(x)=∑kf(k)⌊xk⌋. |
The classical Chebyshev argument shows that
| Ff(x)≤1for every x≥1 | (5) |
implies the rigorous lower bound
| lim infx→∞π(x)x/logx≥A(f) |
[20]. The required global inequality in Equation (5) is much more restrictive than the prime number theorem itself: a single finite weight must satisfy the inequality for every x. AlphaEvolve’s score tested this inequality only at finitely many sampled values. It could therefore assign a high score to a weight that fails at an untested value, in which case the score does not prove the stated prime-counting bound. However, an exhaustive check at all x is usually computationally prohibitive because the associated period can be enormous. The sampled score consequently provides only a rough approximation to whether the global inequality holds.
S1. A score of 0.980681 valid for every x.
The Station agents discovered a finite construction f satisfying Equation (5) for every x, with
| A(f)≥0.980681. | (6) |
This improves on AlphaEvolve’s reported score of 0.938. More importantly, the agents proved the required inequality for all x, whereas the score alone does not provide that guarantee. Their key idea was to choose the integers in the construction so that Ff repeats after a manageable range. This reduces the infinitely many possible values of x to one finite exhaustive check, which the agents completed using exact arithmetic in under a minute.
In contrast, other agents in the same run found constructions with higher scores, reaching 0.990629, but these constructions did not satisfy the global inequality for every x. This provides a concrete example of agents prioritizing the underlying mathematical problem over naive score optimization despite a hackable score.
S2. Why a direct Möbius cutoff fails.
The Möbius function is a natural starting point because it is central to a standard formulation of the prime number theorem. AlphaEvolve explored finite constructions obtained by truncating the Möbius function, and the Station agents initially pursued the same approach. They then proved that this family cannot yield a positive asymptotic score: as the truncation cutoff D grows, its largest violation of the required global inequality grows at least on the order of D/log2D. Consequently, rescaling the construction to satisfy the inequality forces its score down to O(log2D/D), which tends to zero. The proof builds on results about incomplete Möbius sums [45]. This obstruction led the agents to abandon direct Möbius cutoffs and explore a more flexible construction with jointly optimized coefficients, producing the rigorous score of 0.980681 described above.
Limitation.
Since the prime number theorem already determines the limiting ratio above exactly, these results do not change what is known about prime distribution. Their mathematical contribution is narrower: within the finite setting of the benchmark, the Station agents found a construction with a rigorous score of 0.980681 and proved that the natural Möbius cutoff cannot yield a positive asymptotic score. The problem therefore serves primarily as a calibration of whether agents can distinguish a valid mathematical result from a high but hackable score, rather than as a material contribution to the study of prime distribution.
4.9 Difference bases
A finite set B⊂ℤ is a difference basis for {1,…,n} if every integer in that interval is a difference of two elements of B. If Δ(n) is the smallest possible size of such a set, the quantity to minimize is Δ(n)2/n; Rédei and Rényi proved that these normalized minima converge and that their limit is their infimum [65]. AlphaEvolve reported the upper bound
| C:=infn≥1Δ(n)2n≤360249109≈2.639027 |
as Problem 6.7 of its collection. The preceding published upper bound was Golay’s C≤2.6458… [32, 7], rather than the 2.6571… benchmark used in AlphaEvolve’s comparison. This example was found with the help of a human expert hint: the paper records that AlphaEvolve failed to improve its benchmark until it was supplied with correct code for generating Singer difference sets, and its released prompt also directs the search to Singer sets and the classical Leech product construction. We gave the Station only the problem definition, the scoring rule, and a trivial grid baseline. In particular, the agents had neither these construction hints nor access to the external literature.
S1. Independent recovery of a record in the Leech–Golay family.
Leech and Golay combined the four-point difference basis {0,1,4,6} with Singer difference sets to obtain earlier members of this construction family [43, 32, 4]. The Station independently recovered its q=89 member. Taking v=q2+q+1=8011, a 90-element Singer difference set D⊂ℤv, and A={0,1,4,6}, the agents formed
| B={va+d:a∈A,d∈D}. |
With the appropriate representatives for D, the resulting 360 integers realize every difference from 1 through 49109, while 49110 is the first missing difference. Thus
| C≤360249109=2.6390274695…, |
improving Golay’s preceding bound by approximately 0.0067. The set agrees entry for entry with the construction reported by AlphaEvolve. This is an independent recovery of a known record, not a new upper bound relative to AlphaEvolve or a new construction family. The agents also tried to push the lower bound further, but reached only the classical bound C≥2.434467… [43], whereas Yang and Liao proved the stronger published bound C>2.4421 [84].
4.10 Sidorenko’s conjecture
Sidorenko’s conjecture asserts that every bipartite graph H satisfies t(H,W)≥t(K2,W)|E(H)| for every graphon W, where t(H,W) is the homomorphism density of H in W [74]. The smallest unresolved instance is the ten-vertex, fifteen-edge graph H=K5,5∖C10, also called the bipartite Möbius ladder [66]. AlphaEvolve took up this problem as Problem 6.26 of its mathematical collection and searched over nonconstant 30-step graphons. It scored a candidate by
| t(K2,W)15t(H,W)−1, |
so a positive value would give a counterexample and disprove this instance of the conjecture.
AlphaEvolve reported that it did not find a counterexample. We gave the Station the same problem and scoring rule, and the Station agents likewise found none. As such, the status of the conjecture is unchanged.
4.11 Peak autoconvolution
AlphaEvolve’s Problem 6.2, called the first autocorrelation inequality in its collection, asks how evenly the sum of two independent random variables with the same compactly supported density can be distributed. More precisely, for a nonnegative function f supported on [−1/4,1/4] and normalized by ∫f=1, let
| C6.2=inff∥f∗f∥∞. |
Determining C6.2 is connected to the asymptotic size of generalized Sidon sets, and its exact value remains unknown [53]. The best currently reported bounds are
| 1.2937≤C6.2≤1.502851, |
with the lower and upper endpoints coming from certified convex relaxations and an explicit step function, respectively [41, 68].
AlphaEvolve achieved the upper bound C6.2≤1.5032, improving the pre-AlphaEvolve bound C6.2≤1.50972 of Matolcsi and Vinuesa [53]; TTT-Discover later advanced the frontier to C6.2≤1.502863 [86], and an exact-arithmetic certificate improved it further to C6.2≤1.502851 [68]. The Station reached only C6.2≤1.504473, worse than both AlphaEvolve and the current frontier. AlphaEvolve’s highly irregular construction emerged from large-scale heuristic search. This contrast highlights a limitation of the Station: its agents generally favored theory-guided constructions over heuristic search, a preference that produced strong results on several other problems but left them behind here, where frontier constructions depend on extensive heuristic optimization.
4.12 Flat autoconvolution
AlphaEvolve’s Problem 6.3, called the second autocorrelation inequality in its collection, asks how closely the autoconvolution of a nonnegative function can resemble a flat-topped function, constant on a set and zero outside it. More precisely, for a nonzero nonnegative function f∈L1(ℝ)∩L2(ℝ), let
| Q(f)=∥f∗f∥22∥f∗f∥1∥f∗f∥∞,C6.3=supfQ(f). |
Hölder’s inequality gives C6.3≤1; for an arbitrary nonnegative output, equality occurs only for such a flat-topped function. Whether the autoconvolution constraint forces the strict inequality C6.3<1 remains open [51, 53]. Before AlphaEvolve, the best known bounds were [53]
| 0.88922≤C6.3≤1. |
AlphaEvolve established the lower bound C6.3≥0.961021, while later work further improved this to 0.962694 [85]. The Station’s best verified construction reached only C6.3>0.953189 and therefore did not improve the numerical bound. This shortfall reflects the same limitation seen in Problem 6.2, minimizing the peak of an autoconvolution (Section 4.11): the Station’s theory-guided agents were poorly suited to finding the highly irregular constructions produced by large-scale heuristic search.
S1. Binary step functions preserve the unrestricted supremum.
The agents nevertheless proved a useful fact about the search for near-optimal constructions: the supremum defining C6.3 can be approached using binary step functions, thus replacing the search over arbitrary nonnegative functions with a search over binary functions on increasingly fine grids.
4.13 Book Ramsey numbers
Given graphs G1,G2, the Ramsey number R(G1,G2) is the smallest n such that every red-blue edge coloring of Kn forces either a red copy of G1 or a blue copy of G2. Establishing the exact values of Ramsey numbers is a difficult computational and theoretical challenge. The most famous Ramsey numbers are those where G1 and G2 are complete graphs, but many other choices have been studied extensively (see the survey [63]). The book graph Bk consists of k triangles that share a common edge. An open problem is whether
| R(Bn−1,Bn)=4n−1 | (7) |
holds for every positive integer n. Rousseau and Sheehan established the upper bound in 1978, proving R(Bn−1,Bn)≤4n−1 for all n [67]. It therefore remains to prove the matching lower bound. For a given n, this amounts to constructing a red–blue edge coloring of K4n−2 containing neither a red Bn−1 nor a blue Bn.
The third author proved equality for n≤20, independently matching contemporaneous work, and established an infinite Paley-type family whenever 2n−1 is a prime power congruent to 1(mod4) [82, 47]. This combination of finite evidence and a general arithmetic construction led him to conjecture that (7) holds for all n [82]. Epoch AI subsequently adopted it as a FrontierMath open problem [22]. After its posting, further work extended the consecutively solved range to n≤56 and produced two additional infinite families by extending established constructions [80].
We ran two Stations on this problem. The first operated without internet access and discovered a novel conference-graph family. We then ran a second Station with internet access and a summary of the first Station’s results; it discovered a new doubled Legendre family together with several new finite constructions. An external expert subsequently combined the pattern in these finite constructions with an earlier result from the second Station to obtain the Yamada–Pott infinite family. Thus, the first two families are autonomous Station discoveries, whereas the third required human expert involvement. All three families are novel relative to the existing literature and are visualized in Figure 6. The parameters n covered by each family, including which were previously open, are summarized in Figure 7.
S1. A conference-graph family.
The first and broadest family converts any conference graph into a sharp book-Ramsey coloring. Specifically, if a strongly regular graph with parameters
| (q,q−12,q−54,q−14) |
exists, then the Station’s agents proved
| R(Bq,Bq+1)=4q+3. | (8) |
Paley conference graphs exist whenever q is a prime power congruent to 1(mod4). Consequently, the theorem proves the conjecture whenever n−1 is a prime power congruent to 1(mod4). Beyond the Paley case, Seberry and Whiteman used Mathon’s construction to obtain symmetric conference matrices of order 5⋅92t+1+1 for every t≥0 [52, 72]. These yield conference graphs of order q=5⋅92t+1, so the Station theorem also proves the conjecture whenever
| n=5⋅92t+1+1,t≥0. |
The first member gives q=45 and n=46. The known conference graph of order q=65 supplies the additional parameter n=66 [36]. In total, known conference graphs prove the conjecture at 30 values of n≤200, including 19 that were previously open [82, 47, 80, 22].
S2. A doubled Legendre family.
The second family converts a periodic Legendre source over 𝔽Q into a sharp book-Ramsey coloring [25]. Specifically, for every prime power Q>3 with Q≡3(mod8), the Station’s agents proved
| R(B(Q−1)/2,B(Q+1)/2)=2Q+1. | (9) |
Consequently, the theorem proves the conjecture whenever 2n−1 is a prime power congruent to 3(mod8). For n≤200, this family proves equality at 21 values and, at the time of its discovery, resolved six additional open cases after accounting for the conference family [82, 47, 80, 22].
The agents discovered this general family in mid-July 2026. Concurrent work announced at the end of July independently produced the finite case n=70 [22]; the Station theorem contains n=70 as one member and covers infinitely many further parameters.
The doubled Legendre family is related to, but distinct from, the Legendre family reported by Turturean [80]. Both begin with the same type of periodic Legendre source over 𝔽Q, with Q≡3(mod8), but use different lifts to obtain a book-Ramsey coloring. For the same source order Q, the earlier lift reaches n=(Q+1)/4, whereas the Station lift reaches n=(Q+1)/2. It therefore doubles the Ramsey parameter and covers a different set of values, as Figure 7 shows.
S3. A Yamada–Pott family.
The third family converts a classical Yamada–Pott design into a sharp book-Ramsey coloring [3]. Specifically, for every prime power q≥7 with q≡3(mod4), we proved
| R(B(q2−q−2)/4,B(q2−q+2)/4)=q2−q+1. | (10) |
Consequently, the theorem proves the conjecture whenever
| n=q2−q+24 |
for a prime power q≥7 congruent to 3(mod4). For n≤200, this family proves equality at five values and resolves three additional previously open cases after accounting for the conference and doubled Legendre families [82, 47, 80, 22]. The second Station’s agents supplied finite affine constructions for n=11,28,86 and an earlier periodic-correlation identity; an external expert recognized their shared Yamada–Pott structure and used these ingredients to establish the general theorem.
Discussion.
The three infinite families above are novel relative to the existing literature, but their source objects are not: conference graphs, periodic Legendre pairs, and Yamada–Pott designs were all established previously [52, 25, 3]. What is new in each case is the rule that lifts the classical object to a sharp book-Ramsey coloring, and such a rule need not be apparent from the source alone. For example, the agents discovered the general conference-graph lift only after more than 3,000 Station ticks and a long sequence of intermediate internal papers. The accompanying notebook provides the relatively unpolished proofs adapted from the agents’ internal papers; we will present polished proofs of all three families in a separate follow-up paper.
The first two families also show that Station agents can advance a general mathematical objective beyond the directly scorable task: they discovered and proved infinite families even though the evaluator could reward only finite constructions. The third family illustrates a complementary limitation. Both the finite affine examples and the periodic-correlation identity needed for the general theorem were already present in the Station’s research history, but the agents did not connect them. An external expert recognized their shared Yamada–Pott structure and completed the synthesis. This missed connection indicates that agents may not yet capitalize fully on knowledge accumulated across the Station and may benefit from external expert synthesis in such cases.
4.14 Jacobian Conjecture
The Jacobian conjecture asked whether a polynomial map that is locally invertible everywhere must also be globally invertible. More precisely, it asserted that every polynomial map F:ℂn→ℂn with nonzero constant Jacobian determinant is a polynomial automorphism [40]. On 19 July 2026, it was announced that a three-dimensional counterexample had been produced with Claude Fable [1], thereby disproving the conjecture in every dimension at least three. The breakthrough then prompted researchers to seek a conceptual explanation for the map: in particular, why its apparently miraculous Jacobian cancellation occurs and how three generic inverse sheets can coexist with local invertibility everywhere [15, 27, 78, 73, 75].
We launched the Station one week after the announcement. Because this experiment was conducted after newer models had become available, it used a more recent agent pool than the other Stations: two agents each powered by GPT-5.6 Sol, Claude Opus 5, and Gemini 3.1 Pro. The agents had no external web access and received only a formula-free specification: construct a rational-coefficient polynomial map ℂ3→ℂ3 of degree at most 12 with nonzero constant Jacobian determinant and two distinct rational points in one fiber. The evaluator automatically checked each construction and assigned a score of 1 only if it satisfied every requirement, and 0 otherwise. We supplied no literature survey or partial construction. The agents therefore had to find the counterexample independently.
The goal of this task was twofold. First, we wanted to test the Station on a strictly binary problem. The evaluator supplied neither partial credit nor graded feedback, so unsuccessful attempts gave the agents no score signal about how to improve; attaining a score of 1 required reconstructing a counterexample to a conjecture that had resisted mathematicians for nearly nine decades [40]. Second, we wanted to observe the complete discovery process rather than only the final construction. We make the entire raw agent dialogue public, whereas the original Fable discovery trajectory has not been released. This record preserves intermediate mathematical ideas that do not appear in the final construction and allows researchers to study the dynamics of AI-led mathematical discovery.
S1. Independent reconstruction through a cuspidal ruling.
Writing b=xy−1, a Station agent constructed the degree-seven map
| F(x,y,z)=(6x+9x2y,y(9b2+6b−2), 3y2b(3b−1))+z(x3,xb2,b3). |
Exact calculation gives detJF=−6, and the three distinct rational points
| (−67,−76,−4753216),(34,73,−98027),(328,73,254827) |
all map to (1,7/3,0). These identities constitute a complete counterexample certificate. The formula differs visibly from the announced map H [1, 26], but the linear source and target transformations T(x,y,z)=(x,−y,−3z) and L(A,B,C)=(3C,−B,3A) satisfy F∘T=L∘H. The Station therefore reconstructed the announced counterexample in different linear coordinates; it did not produce a new counterexample or a new equivalence class.
Whereas the original result was credited to Claude Fable, the counterexample was independently discovered within one day by a single GPT-5.6 Sol agent, without direct interaction with the other agents. The successful agent began with ruled maps F(x,y,z)=f(x,y)+zn(x,y), so that varying z traces a line for each fixed (x,y). It tested five low-degree direction templates based on smooth conics, but none satisfied the remaining constant-Jacobian condition. The decisive step was to replace the smooth direction curve with the cuspidal cubic [r:s]↦[r3:rs2:s3]. Its associated direction field is n=(x3,x(xy−1)2,(xy−1)3); with this choice, the compatibility equations for the base surface f became solvable and yielded exactly the map above.
S2. The reconstructed map has three-sheeted fibers without critical points.
During the successful derivation, the agent also explained why the cuspidal ruling makes the Jacobian constant. For the direction field n=(x3,x(xy−1)2,(xy−1)3), the agent derived moving-frame identities, including D(n)=3xn for D=x2∂x−∂y, under which every z-dependent contribution to the determinant contains a repeated tangent direction and vanishes. The remaining triple product is the constant −6. The agent thus derived the Jacobian cancellation from the geometry of the cuspidal ruling rather than discovering sixteen terms whose cancellation could only be checked afterward.
After constructing the counterexample, the same agent analyzed its fibers and explained how the map can be locally invertible everywhere while generically having three preimages. On a dense chart, write a target as (X,Y,Z) and set I=XY and J=X2Z. Recovering a preimage then reduces to
| t3+6t2−3It+2J=0. | (11) |
For a generic target, the three roots give three distinct preimages. If p(t) denotes the left-hand side, the inverse formulas satisfy A=p′(t)/6, X=xA, and hence x=X/A. When roots coalesce and X≠0, the condition p′(t)=0 forces the corresponding source point to escape to infinity rather than become a critical point in affine space. Over the exceptional locus X=0, the source coordinate x supplies an additional affine scale direction that resolves the same apparent ramification. This analysis answers the structural question raised by mathematicians immediately after the announcement: the three sheets arise from a cubic quotient, while the geometry of the full three-dimensional map prevents their collisions from producing critical points. The agent’s explanation coincides with the cuspidal and cubic account developed by mathematicians in the days following the announcement [27, 78, 73, 75].
Discussion.
The mathematical outcome of this experiment is an independent reconstruction, not a new counterexample or a new explanation. The example indicates that the Station can tackle a difficult binary problem whose evaluator provides no gradient or partial score to guide the search. Counterexample breakthroughs of this kind may nevertheless be rare because conjectures are generally expected to be true. In a broader context, the harder challenge may therefore be identifying a promising problem and investing substantial computation before knowing whether a counterexample exists.
5 Meta-analysis
In this section, we perform a meta-analysis of the discovery process above to better understand the dynamics of AI discovery. Unless otherwise stated, all analyses are based on the 16 Station instances behind the 14 problems mentioned in the preceding section. (The Kissing number in d=11 and Book Ramsey numbers problems each have two Station instances.) Spotlight results refer to the results marked S1, S2, and so forth in that section, totaling 28 results. When a single spotlight contains multiple independently discovered findings, we count those findings separately. We use archive paper to refer to a paper published by an agent within the Station, not a paper in the external human literature.
5.1 Contributions from model families
We first analyze the primary contributor to each of the 28 spotlight results, as shown in Figure 8(a). We attribute each result to the agent that made the substantive discovery, rather than to an agent that later restated, verified, or published it. Claude agents made the primary discovery for 18 results (64.3%), GPT agents for 9 (32.1%), and Gemini agents for 1 (3.6%). Gemini’s smaller share may partly reflect model ages: Gemini 3.1 Pro was released in February 2026, earlier than GPT-5.5 in April and Claude Opus 4.8 in May [34, 60, 2]. Its lower contribution is therefore consistent with the general industry trend of later model releases achieving stronger capabilities.
We also analyze the agents’ archive paper contributions, as shown in Figure 8(b). Gemini agents submitted the most archive papers: 2,652 attempts, of which 508 were accepted (19.2%), so more than 80% were rejected by the reviewer. Claude agents made 1,236 attempts, of which 696 were accepted (56.3%), while GPT agents made only 506 attempts, of which 388 were accepted (76.7%). We also compute the total citations by model family and find that archive papers by Claude agents received the most citations both in total and on average (Figure 8(c)). In our observation, Gemini agents tended to overclaim, for example by declaring a direction impossible on the basis of limited evidence; such submissions were generally rejected by the reviewer system, which may help explain the high rejection rate. In contrast, GPT agents were very prudent in archive paper submission and often submitted only when a finding was relatively material, which may help explain the low submission count. Claude archive papers were generally much longer and more comprehensive, which may help explain their higher average citation count. These patterns reflect the different research styles of the model families.
Qualitatively, we observe substantial differences in the strengths and failure modes of the three model families. Gemini agents tended to propose more novel heuristics and research directions, but they were also more likely to overstate claims or change course too readily in response to peer feedback. GPT agents tended to be more rigorous and were often able to produce valid informal proofs of new results, but they could become absorbed in technically intricate side questions whose broader research value was limited. Claude agents tended to be persistent, methodical, and self-critical. Their creativity was often adaptive: they learned from failed approaches, used those failures to identify new directions, and pursued those directions persistently through rigorous verification. This combination of rigor and disciplined creativity made Claude a prolific contributor. Its agents nevertheless occasionally made erroneous claims that were later corrected by peer agents.
5.2 Collaboration across model families
One characteristic of the Station is that it allows agents from different model families to collaborate. We therefore ask how often agents from different model families worked together on a spotlight result. We examine all 28 spotlight results above. We count an agent as a contributor when its work was used materially in the result, for example when it contributed a theorem, construction, method, or research direction that another agent used.
We find that 13 of the 28 spotlight results (46.4%) involved agents from more than one model family, as shown in Figure 9(a). Among the remaining 15 results, 6 were still joint work by several agents from the same model family. Thus, only 9 of the 28 results (32.1%) were found by one agent working alone, while 19 (67.9%) involved more than one agent. Claude agents were particularly collaborative: they took part in all 13 cross-model results. These findings suggest that collaboration across agents and model families was an important part of the discovery process. Most current AI-for-science systems, by contrast, either use agents from a single model family within a run [58, 30, 70, 35, 61], or use different model families in fixed roles within a pipeline [50, 31].
We also examine how agents communicated in the cross-model cases. The Archive Room was the most frequent channel, accounting for 61.5% of these collaborations (Figure 9(b)). This suggests that archive papers are an efficient means of peer communication. As highly distilled accounts of scientific outcomes from an agent’s longer research process, these archive papers provide a low-bandwidth but information-dense body of knowledge on which later agents can build, much like our own scientific literature. One agent could solve part of a problem and explain what was still missing; a later agent from another model family could read the archive paper and continue. Indeed, a prominent case study of collaboration among three model families, conducted mostly through archive papers and leading to the first finite-Kakeya spotlight result, is shown in Figure 9(c).
5.3 Discovery time
We are also interested in how long the Station took to make each discovery. Most Station instances ran for 1,000–2,000 ticks, corresponding to roughly one to two weeks of continuous wall-clock time. Figure 10 shows the tick at which each of the 28 spotlight results first appeared in its final substantive form.
Some relatively simple results appeared early. With the notable exception of the Jacobian Conjecture, these early discoveries tended to be less substantial, often consisting of relatively direct adaptations or extensions of ideas available from pretrained knowledge, before much shared Station knowledge had accumulated.
Thirteen of the 28 spotlight results (46.4%) were discovered after tick 1000. We generally observed that later discoveries tended to be more novel or difficult. The most extreme example was the conference-graph family for Book Ramsey numbers, discovered at tick 3727. Its lifting rule was far from obvious from the existing literature and warranted a separate external follow-up paper. Such nontrivial discoveries often emerged only after a substantial internal literature had accumulated.
5.4 Station mechanisms
The Station is designed to foster scientific discovery through several mechanisms. These mechanisms are described in detail in Appendix A; here we give a brief overview and ask which of them contributed to the spotlight results.
Holiday. The final two ticks of every ten-tick period are declared a holiday; agents cannot submit code or archive papers and instead receive prompts encouraging broad reflection, metaphors, or ideas from other fields. This pause often led agents to reconsider a failed approach or explore a less obvious direction.
Archive paper. Accepted archive papers form the Station’s cumulative knowledge and remain available to later agents. This allows partial theorems, constructions, and well-documented failures to become starting points for later discoveries.
Stagnation protocol. If the official evaluation frontier does not improve for a long period, the Station asks agents to review the internal literature, question their assumptions, and pursue different high-level strategies. This helps agents leave exhausted local approaches and pushes them toward bolder attempts and wider exploration.
Peer communication. Agents can exchange partial results, targeted questions, and criticism through direct mail or shared public discussion.
Supervisor. The Station randomly appoints one eligible agent to serve as supervisor. The supervisor gives high-level guidance, encouraging persistence and preventing agents from duplicating one another’s work while leaving them responsible for their own research; between appointments, the Station deliberately leaves long periods without a supervisor to encourage less structured exploration.
Question Room. Agents can post important open subproblems for other agents to discuss and solve. This turns unresolved gaps into shared research targets and allows agents with different approaches to supply missing pieces.
These mechanisms support discovery in different ways. Holidays widen exploration; archive papers deepen cumulative knowledge; the stagnation protocol provides a push away from local optima; and peer communication, supervision, and the Question Room coordinate work across agents.
We reviewed the dialogue underlying each of the 28 results and classified each mechanism as making a direct contribution, an indirect contribution, or no material contribution to the discovery (Figure 11). A contribution was direct when the mechanism supplied a decisive idea or intervention, and indirect when it shaped or supported the research without being the immediate source of the result. We assigned no material contribution when the dialogue showed no clear causal role.
Holiday and archive papers contributed directly or indirectly to 23 and 21 of the 28 results, respectively, followed by the stagnation protocol with 14. During holidays, agents often stepped back from active optimization, examined why an earlier approach had failed, and reframed the problem or explored a new direction; these reflections frequently supplied ideas that later became part of a spotlight result, explaining the high contribution rate. Archive papers also contributed to a significant portion of the results, indicating that the Station’s accumulated knowledge was useful for later discoveries.
5.5 Result reproducibility
We are also interested in whether the discoveries are reproducible. We therefore ran three independent Station instances, all without web access, on the kissing-number problem in dimension eleven. (These include the two instances described in Section 4.3; the third is used only for this reproducibility analysis and is not included in the other meta-analyses above.) Figure 12 shows the best certified lower bound reached in each run. All three Stations eventually reached N=604, indicating that the improved lower bound is reproducible.
Closer inspection, however, shows substantial variation in both the time required and the route to the result. Station 1 pursued discrete exact line packing around lattice-derived cores. It obtained Construction 3 by selecting 54 mutually compatible lines that form a 108-point algebraic extension of a 496-point core, and later obtained Construction 2 while exploring a different core and extension. Station 2 instead assembled Construction 1 from root-system motifs under a common rotation; its final step was to recognize that eleven points formed all but one vertex of a cuboctahedron and to add the missing twelfth vertex. Station 3 reached the same construction class through a different mechanism: it deformed an exact 601-point configuration so that two coordinate vectors and one additional vector supported on a distinguished three-dimensional subspace could be appended. Thus, the same numerical lower bound emerged from markedly different mathematical representations and research paths.
This variation partly arises from the Station’s cumulative knowledge. Small differences in the initial trajectory change which results enter the archive paper collection. Later agents then inherit different starting points, so differences in research paths and accumulated archive papers compound over time. Therefore, given the high variance across Station instances, running several independent instances on the same problem is advisable when computational cost is not a concern.
6 Discussion and Conclusion
We observe rapid improvement in the capabilities of AI agents. In the initial version one year ago, agents frequently hallucinated and could not reliably learn the rules of the environment. Agents can now master the environment and autonomously produce novel discoveries. Nonetheless, multi-agent research still has several important limitations. We summarize our observations below.
Lack of expert intuition. By intuition, we mean the ability to judge whether a research direction is promising before pursuing it. Good intuition makes exploration more efficient and allows a researcher to investigate promising directions more deeply. Across the runs, we observed multiple cases in which agents deprioritized promising approaches on weak grounds, delaying or missing potential breakthroughs. This indicates a lack of the intuition that a human expert in the field would typically possess.
Lack of diverse research tastes. A preference for particular concepts or methods is difficult to judge as objectively good or poor. However, when all agents share similar tastes, the overall scope of exploration becomes narrow. Across the runs, agents from the same model family often proposed similar research ideas, suggesting that model-specific tastes reduce the diversity of exploration.
Limited in-context learning. Agents can absorb new research knowledge through their context, but this knowledge does not update their pretrained weights. As the Station’s accumulated knowledge grows, agents may therefore struggle to absorb it fully and build on it effectively. We occasionally observed agents fail to recognize how their own line of research connected to earlier Station knowledge, causing them to miss a potential discovery.
Attractor traps. When given autonomy, some agents become absorbed in tasks or activities that we call attractors. These activities are often rewarding in some immediate sense but make little meaningful contribution to the main problem. Agents may also become absorbed in technical details that a human expert would quickly recognize as trivial or irrelevant to the main question. Examples include repeatedly rerunning the same optimization script with different random seeds or exhaustively diagnosing and characterizing every local optimum.
Several Station mechanisms are designed to mitigate these limitations. For example, using agents from multiple model families broadens the range of research tastes, while the stagnation protocol helps agents escape attractor traps. Nevertheless, these problems persist to some degree, and substantial gaps remain between AI agents and human experts in all four respects. Lightweight guidance or occasional intervention from human experts would likely be beneficial by directing agents toward promising research areas. The current Station supports such human involvement, e.g., through messages broadcast to all agents, but we leave a systematic study of human–AI collaboration to future work.
Although this paper uses the Station primarily for mathematical exploration, the Station is designed as a general research environment, and none of its mechanisms is tailored specifically to mathematics. As demonstrated in the original paper, the Station can be applied to problems spanning mathematics, computational biology, and machine learning [16]. Large-scale research explorations in other fields, including research on language models themselves, may therefore be promising.
As AI agents become more capable, we expect autonomy and generality to become increasingly important principles for designing AI research environments. Stronger agents need not be confined to increasingly elaborate pipelines; they have the ability to determine how to pursue a goal, learn from failure, exchange ideas, and accumulate knowledge over time. The greater autonomy provided by the Station may allow these capabilities to be more fully realized.
References
- [1] L. Alpöge (2026) Hello there the Jacobian conjecture is false. Note: X postPosted 19 July 2026 External Links: Link Cited by: §1, §4.14, §4.14.
- [2] Anthropic (2026) Claude opus 4.8. Note: Anthropic External Links: Link Cited by: §5.1.
- [3] K. T. Arasu, D. A. Bulutoglu, and J. R. Hollon (2020) Legendre G-array pairs and the theoretical unification of several G-array families. Journal of Combinatorial Designs 28 (11), pp. 814–841. Note: arXiv:2004.05608 External Links: Document Cited by: §4.13, §4.13.
- [4] T. Banakh and V. Gavrylkiv (2019) Difference bases in cyclic groups. Journal of Algebra and Its Applications 18 (5), pp. 1950081. Note: arXiv:1702.02631 External Links: Document Cited by: §4.9.
- [5] R. D. Benguria and M. Loss (2004) Connection between the Lieb–Thirring conjecture for Schrödinger operators and an isoperimetric problem for ovals on the plane. In Partial Differential Equations and Inverse Problems, Contemporary Mathematics, Vol. 362, pp. 53–61. Note: arXiv:math-ph/0402048 Cited by: §4.7, §4.7.
- [6] A. Bernal (1989) A note on the one-dimensional maximal function. Proceedings of the Royal Society of Edinburgh Section A: Mathematics 111 (3–4), pp. 325–328. External Links: Document Cited by: §4.6.
- [7] A. Bernshteyn and M. Tait (2019) Improved lower bound for difference bases. Journal of Number Theory 205, pp. 50–58. Note: arXiv:1901.09411 External Links: Document Cited by: §4.9.
- [8] J. Bernstein and T. Mettler (2015) One-dimensional projective structures, convex curves and the ovals of Benguria & Loss. Communications in Mathematical Physics 336 (2), pp. 933–952. Note: arXiv:1403.8000 External Links: Document Cited by: §4.7, §4.7.
- [9] M. R. Best (1977) A(11,4,4)=35, Or some new optimal constant-weight codes. Technical report Technical Report ZN 71/77, Mathematical Centre, Amsterdam. External Links: Link Cited by: §4.3.
- [10] F. Bianchi, Y. Kwon, A. Pappu, and J. Zou (2026) Harnessing the collective intelligence of AI agents in the wild for new discoveries. arXiv preprint arXiv:2606.10402. External Links: Document Cited by: §4.3.
- [11] A. Blokhuis and F. Mazzocca (2008) The finite field kakeya problem. In Building Bridges: Between Mathematics and Computer Science, M. Grötschel and G. O. H. Katona (Eds.), Bolyai Society Mathematical Studies, Vol. 19, pp. 205–218. Note: arXiv:0911.4370 External Links: Document Cited by: 4th item.
- [12] J. Bourgain, L. Clozel, and J. Kahane (2010) Principe d’Heisenberg et fonctions positives. Annales de l’Institut Fourier 60 (4), pp. 1215–1232. External Links: Document Cited by: §4.5.
- [13] B. Bukh and T. Chao (2021) Sharp density bounds on the finite field kakeya problem. Discrete Analysis. Note: Article 26, 9 pp.; arXiv:2108.00074 External Links: Document Cited by: 1st item, §4.1, §4.1, §4.1.
- [14] A. Burchard and L. E. Thomas (2005) On an isoperimetric inequality for a Schrödinger operator depending on the curvature of a loop. The Journal of Geometric Analysis 15 (4), pp. 543–563. Note: arXiv:math/0505123 External Links: Document Cited by: §4.7, §4.7.
- [15] K. Buzzard (2026) Human mathematicians are being outcounterexampled. Note: The Xena Project blog External Links: Link Cited by: §4.14.
- [16] S. Chung and W. Du (2025) The station: an open-world environment for ai-driven discovery. External Links: 2511.06309, Document, Link Cited by: Appendix A, §1, §2, §6.
- [17] H. Cohn and F. Gonçalves (2019) An optimal uncertainty principle in twelve dimensions via modular forms. Inventiones Mathematicae 217, pp. 799–831. Note: arXiv:1712.04438 External Links: Document Cited by: §4.5.
- [18] H. Cohn (2026) Kissing numbers. Note: Online tablehttps://cohn.mit.edu/kissing-numbers/, accessed 4 August 2026 Cited by: §4.3, §4.3.
- [19] A. Córdoba (1977) The kakeya maximal function and the spherical summation multipliers. American Journal of Mathematics 99 (1), pp. 1–22. External Links: Document Cited by: §4.4.
- [20] H. G. Diamond (1982) Elementary methods in the study of the distribution of prime numbers. Bulletin of the American Mathematical Society 7 (3), pp. 553–589. External Links: Document Cited by: §4.8.
- [21] Z. Dvir (2009) On the size of kakeya sets in finite fields. Journal of the American Mathematical Society 22 (4), pp. 1093–1097. External Links: Document Cited by: §4.1.
- [22] Epoch AI (2026) Book Ramsey numbers. Note: FrontierMath Open ProblemsAccessed 17 August 2026 External Links: Link Cited by: §4.13, §4.13, §4.13, §4.13, §4.13.
- [23] P. Erdős (1955) Some remarks on number theory. Riveon Lematematika 9, pp. 45–48. Note: In Hebrew Cited by: §4.2.
- [24] K. J. Falconer (1985) The geometry of fractal sets. Cambridge Tracts in Mathematics, Vol. 85, Cambridge University Press. Cited by: §4.4, §4.4.
- [25] R. J. Fletcher, M. Gysin, and J. Seberry (2001) Application of the discrete Fourier transform to the search for generalised Legendre pairs and Hadamard matrices. Australasian Journal of Combinatorics 23, pp. 75–86. External Links: Link Cited by: §4.13, §4.13.
- [26] A. Freitas Ramos, D. Barros Hulak, and R. J. Guerra Barretto de Queiroz (2026) Formal verification of an explicit counterexample to the Jacobian conjecture. Note: Archive of Formal Proofs External Links: Link Cited by: §4.14.
- [27] A. Gallagher (2026) An infinite family of counterexamples to the Jacobian conjecture in dimension three: every generic fiber degree n≥3 occurs. Note: Zenodo preprint External Links: Document, Link Cited by: §4.14, §4.14.
- [28] M. Ganzhinov (2025) Highly symmetric lines. Linear Algebra and its Applications 722, pp. 12–37. Note: arXiv:2207.08266 External Links: Document Cited by: §4.3.
- [29] B. Georgiev, J. Gómez-Serrano, T. Tao, and A. Z. Wagner (2025) Mathematical exploration and discovery at scale. arXiv preprint arXiv:2511.02864. External Links: Document, Link Cited by: §3.1, §3.2.
- [30] A. Ghafarollahi and M. J. Buehler (2025) SciAgents: automating scientific discovery through bioinspired multi-agent intelligent graph reasoning. Advanced Materials 37 (22), pp. 2413523. External Links: Document, Link Cited by: §2, §5.2.
- [31] A. E. Ghareeb, B. Chang, L. Mitchener, A. Yiu, C. J. Szostkiewicz, D. Shved, G. J. Gyimesi, J. M. Laurent, S. M. Wright, M. T. Razzak, A. D. White, S. C. Finnemann, M. M. Hinks, and S. G. Rodriques (2026) A multi-agent system for automating scientific discovery. Nature 655, pp. 497–505. External Links: Document, Link Cited by: §2, §5.2.
- [32] M. J. E. Golay (1972) Notes on the representation of 1,2,…,n by differences. Journal of the London Mathematical Society s2-4 (4), pp. 729–734. External Links: Document Cited by: §4.9, §4.9.
- [33] F. Gonçalves, D. Oliveira e Silva, and S. Steinerberger (2017) Hermite polynomials, linear flows on the torus, and an uncertainty principle for roots. Journal of Mathematical Analysis and Applications 451 (2), pp. 678–711. External Links: Document Cited by: §4.5.
- [34] Google (2026) Introducing Gemini 3.1 Pro: a smarter model for your most complex tasks. Note: Google blog External Links: Link Cited by: §5.1.
- [35] J. Gottweis, W. Weng, A. Daryin, T. Tu, P. Sirkovic, A. Myaskovsky, G. Glowaty, F. Weissenberger, A. Orlandi, D. Popovici, A. Palepu, K. Rong, R. Tanno, K. Saab, F. Zhang, J. Blum, A. Carroll, K. Kulkarni, N. Tomašev, D. Zverinski, I. Rendulic, E. Vedadi, F. Hasler, L. Rimanic, M. Boia, I. Budiselic, B. Feinstein, M. Bellaiche, T. Sheffer, J. Freyberg, J. Ratcliff, O. Bertolli, K. Chou, A. Hassidim, B. Gokturk, A. Vahdat, Y. Guan, V. Dhillon, E. D. Vaishnav, B. Lee, T. R. D. Costa, J. R. Penadés, G. Peltz, Y. Matias, J. Manyika, D. Hassabis, Y. Xu, P. Kohli, A. Pawlosky, A. Karthikesalingam, and V. Natarajan (2026) Accelerating scientific discovery with Co-Scientist. Nature 655, pp. 487–496. External Links: Document, Link Cited by: §2, §5.2.
- [36] O. Gritsenko (2021) On strongly regular graph with parameters (65,32,15,16). arXiv preprint arXiv:2102.05432. External Links: Document Cited by: §4.13.
- [37] J. K. Haugland (2016) The minimum overlap problem revisited. arXiv preprint arXiv:1609.08000. External Links: Document Cited by: §4.2.
- [38] E. Hedley (2025) Can creativity in science be learnt? these researchers think so. Nature. External Links: Document Cited by: §A.4.
- [39] U. Keich (1999) On Lp bounds for kakeya maximal functions and the minkowski dimension in ℝ2. Bulletin of the London Mathematical Society 31 (2), pp. 213–221. External Links: Document Cited by: §4.4.
- [40] O. Keller (1939) Ganze Cremona-transformationen. Monatshefte für Mathematik und Physik 47, pp. 299–306. External Links: Document Cited by: §4.14, §4.14.
- [41] S. Kim and M. Pilanci (2026) AI-assisted discovery of convex relaxations via dual agents. arXiv preprint arXiv:2606.31182. External Links: Document Cited by: §4.11, §4.2, §4.2.
- [42] S. Kopparty, V. F. Lev, S. Saraf, and M. Sudan (2011) Kakeya-type sets in finite vector spaces. Journal of Algebraic Combinatorics 34 (3), pp. 337–355. Note: arXiv:1003.3736 External Links: Document Cited by: 3rd item.
- [43] J. Leech (1956) On the representation of 1,2,…,n by differences. Journal of the London Mathematical Society s1-31 (2), pp. 160–169. External Links: Document Cited by: §4.9, §4.9.
- [44] Leiden Declaration Working Group (2026) Leiden declaration on artificial intelligence and mathematics. External Links: Document, Link Cited by: §1.
- [45] P. Letendre (2020) Truncated convolution of the möbius function and multiplicative energy of an integer n. Acta Arithmetica 195 (1), pp. 83–95. External Links: Document Cited by: §4.8.
- [46] V. F. Lev (2009) Comment 994 on “DHJ3: 900–999 (density Hales–Jewett type numbers)”. Note: Blog comment, What’s new (T. Tao)https://terrytao.wordpress.com/2009/03/04/dhj3-900-999-density-hales-jewett-type-numbers/comment-page-3/#comment-36694, accessed 30 July 2026 Cited by: 6th item, §4.1.
- [47] B. Lidický, G. McKinley, F. Pfender, and S. Van Overberghe (2025) Small Ramsey numbers for books, wheels, and generalizations. The Electronic Journal of Combinatorics 32 (4), pp. P4.64. Note: arXiv:2407.07285 External Links: Document Cited by: Figure 7, Figure 7, §4.13, §4.13, §4.13, §4.13.
- [48] H. Linde (2025) An improved bound for the ground state of a Schrödinger operator on a loop. arXiv preprint arXiv:2504.20229. External Links: Document Cited by: §4.7.
- [49] E. Lorist and F. L. Schwenninger (2026) A solution to Crouzeix’s conjecture. arXiv preprint arXiv:2608.03841. External Links: Document, Link Cited by: §1.
- [50] C. Lu, C. Lu, R. T. Lange, Y. Yamada, S. Hu, J. Foerster, D. Ha, and J. Clune (2026) Towards end-to-end automation of AI research. Nature 651, pp. 914–919. External Links: Document, Link Cited by: §2, §5.2.
- [51] G. Martin and K. O’Bryant (2009) The supremum of autoconvolutions, with applications to additive number theory. Illinois Journal of Mathematics 53 (1), pp. 219–235. External Links: Document Cited by: §4.12.
- [52] R. Mathon (1978) Symmetric conference matrices of order pq2+1. Canadian Journal of Mathematics 30 (2), pp. 321–331. External Links: Document Cited by: §4.13, §4.13.
- [53] M. Matolcsi and C. Vinuesa (2010) Improved bounds on the supremum of autoconvolutions. Journal of Mathematical Analysis and Applications 372 (2), pp. 439–447. External Links: Document Cited by: §4.11, §4.11, §4.12.
- [54] L. Mazur (2026) A computer-assisted proof of Sendov’s conjecture. Note: Proof Atlas External Links: Link Cited by: §1.
- [55] A. D. Melas (2002) On the centered Hardy–Littlewood maximal operator. Transactions of the American Mathematical Society 354, pp. 3263–3273. External Links: Document Cited by: §4.6.
- [56] A. D. Melas (2003) The best constant for the centered Hardy–Littlewood maximal inequality. Annals of Mathematics 157 (2), pp. 647–688. External Links: Document Cited by: §4.6, §4.6.
- [57] G. Mockenhaupt and T. Tao (2004) Restriction and kakeya phenomena for finite fields. Duke Mathematical Journal 121 (1), pp. 35–74. External Links: Document Cited by: 2nd item.
- [58] A. Novikov, N. Vũ, M. Eisenberger, E. Dupont, P. Huang, A. Z. Wagner, S. Shirobokov, B. Kozlovskii, F. J. R. Ruiz, A. Mehrabian, M. P. Kumar, A. See, S. Chaudhuri, G. Holland, A. Davies, S. Nowozin, P. Kohli, and M. Balog (2025) AlphaEvolve: a coding agent for scientific and algorithmic discovery. arXiv preprint arXiv:2506.13131. External Links: Document, Link Cited by: §1, §2, §5.2.
- [59] OpenAI (2026) Codex CLI. Note: OpenAI documentation External Links: Link Cited by: §A.3.
- [60] OpenAI (2026) Introducing GPT-5.5. Note: OpenAI External Links: Link Cited by: §5.1.
- [61] OpenAI (2026) Multi-agent. Note: Accessed: 2026-08-14 External Links: Link Cited by: §2, §5.2.
- [62] OpenAI (2026) Ten advances in mathematics and theoretical computer science. Note: OpenAI External Links: Link Cited by: §1.
- [63] S. P. Radziszowski (2026) Small Ramsey numbers. Electronic Journal of Combinatorics. Note: Dynamic Surveys, DS1, version 18, 24 April 2026 External Links: Document Cited by: §4.13.
- [64] J. P. G. Ramos (2019) Sharp total variation results for maximal functions. Annales Academiae Scientiarum Fennicae Mathematica 44 (1), pp. 41–64. External Links: Document Cited by: §4.6.
- [65] L. Rédei and A. Rényi (1949) On the representation of the numbers 1,2,…,N by means of differences. Matematicheskii Sbornik, New Series 24(66) (3), pp. 385–389. Note: In Russian External Links: Link Cited by: §4.9.
- [66] B. Rossman (2025) On Sidorenko’s conjecture for bipartite Möbius ladders. Note: Preprint External Links: Link Cited by: §4.10.
- [67] C. C. Rousseau and J. Sheehan (1978) On Ramsey numbers for books. Journal of Graph Theory 2 (1), pp. 77–87. External Links: Document Cited by: §4.13.
- [68] K. Russell (2026) Exact-arithmetic certificates for three autoconvolution inequalities, with machine-verified re-evaluations of four published constructions. Zenodo. External Links: Document Cited by: §4.11, §4.11.
- [69] S. Saraf and M. Sudan (2008) An improved lower bound on the size of kakeya sets over finite fields. Analysis & PDE 1 (3), pp. 375–379. Note: arXiv:0808.2499 External Links: Document Cited by: 2nd item.
- [70] S. Schmidgall, Y. Su, Z. Wang, X. Sun, J. Wu, X. Yu, J. Liu, M. Moor, Z. Liu, and E. Barsoum (2025) Agent laboratory: using LLM agents as research assistants. In Findings of the Association for Computational Linguistics: EMNLP 2025, Suzhou, China, pp. 5977–6043. External Links: Document, Link Cited by: §2, §5.2.
- [71] I. J. Schoenberg (1962) On certain minima related to the Besicovitch–Kakeya problem. Mathematica (Cluj) 4, pp. 145–148. Cited by: §4.4.
- [72] J. Seberry and A. L. Whiteman (1988) New Hadamard matrices and conference matrices obtained via Mathon’s construction. Graphs and Combinatorics 4, pp. 355–377. External Links: Document Cited by: §4.13.
- [73] T. Shaska (2026) Graded Keller maps and the Jacobian conjecture. arXiv preprint arXiv:2607.20210. External Links: Document, Link Cited by: §4.14, §4.14.
- [74] A. Sidorenko (1993) A correlation inequality for bipartite graphs. Graphs and Combinatorics 9, pp. 201–204. External Links: Document Cited by: §4.10.
- [75] D. E. Speyer (2026) The geometry and structure of Gallagher’s counterexamples to the Jacobian conjecture. External Links: Link Cited by: §4.14, §4.14.
- [76] R. Takhanov, Z. Assylbekov, and S. Yun (2026) Structure of kissing arrangements in ℝ12 and a place for the 841st sphere. arXiv preprint arXiv:2606.18984. External Links: Document Cited by: §4.3, §4.3.
- [77] R. Takhanov and S. Yun (2026) Classification of independent sets in signed Johnson graphs and applications to kissing arrangements. arXiv preprint arXiv:2606.03299. External Links: Document Cited by: §4.3, §4.3.
- [78] T. Tao (2026) A digestion of the Jacobian conjecture counterexample. Note: What’s New External Links: Link Cited by: §4.14, §4.14.
- [79] T. Tao (2026) Mathematics in the age of AI. arXiv preprint arXiv:2608.16753. External Links: Document, Link Cited by: §1.
- [80] D. Turturean (2026) Summary of new results on the Ramsey numbers for book graphs open problem. Note: Public progress report External Links: Link Cited by: Figure 7, Figure 7, §4.13, §4.13, §4.13, §4.13, §4.13.
- [81] E. Y. Wang, S. Motwani, J. V. Roggeveen, E. Hodges, D. Jayalath, C. London, K. Ramakrishnan, F. Cipcigan, P. Torr, and A. Abate (2026) HorizonMath: measuring AI progress toward mathematical discovery with automatic verification. arXiv preprint arXiv:2603.15617. Cited by: §4.4.
- [82] W. J. Wesley (2026) Lower bounds for book Ramsey numbers. Discrete Mathematics 349, pp. 114913. Note: arXiv:2410.03625 External Links: Document Cited by: Figure 7, Figure 7, §4.13, §4.13, §4.13, §4.13.
- [83] E. P. White (2023) A new bound for Erdős’ minimum overlap problem. Acta Arithmetica 208 (3), pp. 235–255. External Links: Document Cited by: §4.2, §4.2.
- [84] S. Yang and Q. Liao (2022) The lower bound for difference bases. Scientia Sinica Mathematica 52 (11), pp. 1237–1254. Note: In Chinese External Links: Document Cited by: §4.9.
- [85] H. Ye, H. Lin, J. Tang, Y. Luo, R. Thapa, C. Yang, C. Su, R. Yang, R. Liu, R. Li, Z. Li, P. Sun, C. Gao, D. Ding, G. He, M. Zhang, L. Sun, W. Wang, Y. Zhong, Z. Shen, P. Li, P. Lu, B. Cui, D. He, J. Ma, J. Li, H. Baoyin, Y. Choi, S. Ermon, X. Chu, T. Li, Y. Xu, and J. Zou (2026) Structured scaling of AI discovery across diverse scientific domains. arXiv preprint arXiv:2604.19341. External Links: Document Cited by: §4.12, §4.2, §4.2.
- [86] M. Yuksekgonul, D. Koceja, X. Li, F. Bianchi, J. McCaleb, X. Wang, J. Kautz, Y. Choi, J. Zou, C. Guestrin, and Y. Sun (2026) Learning to discover at test time. arXiv preprint arXiv:2601.16175. External Links: Document Cited by: §4.11.
- [87] V. A. Zinoviev and T. Ericson (1999) New lower bounds for contact numbers in small dimensions. Problems of Information Transmission 35 (4), pp. 287–294. External Links: Link Cited by: §4.3.
Appendix A The Station
This appendix provides a self-contained description of the Station used in this paper, which we call Station v2 to distinguish it from the original Station v1. We focus on its mechanisms and implementation details, and refer readers to the original Station paper for the broader design philosophy and motivation behind the environment [16]. The source code is available at https://github.com/dualverse-ai/station.
A.1 Space, Time, and Action
Space.
The Station is divided into rooms, each serving a different purpose (Table 1). For example, agents conduct experiments in the Research Center, read and publish papers in the Archive Room, and communicate with peers in the Mail Room. An agent must be present in a room to use its actions and can move between rooms through navigation actions. This division into rooms gives the environment a modular design with a clear separation of functions.
Time.
The Station operates in discrete time steps called ticks. A tick is completed after every active agent has received one Station observation and returned one response. Ticks provide a shared timeline for all agents in the Station.
In Station v1, agents received their observations sequentially. In contrast, Station v2 first prepares an observation for every agent from the same state at the beginning of the tick and then sends the observations to all agents in parallel. This substantially reduces the wall-clock time required for a Station run.
Action.
At each tick, an agent receives an observation containing its current status, new system messages, the outcomes of its previous actions, and the latest output from the rooms it visited. The agent replies with free-form text together with any actions it intends to perform. Actions are written using the command /execute_action{...} and may be followed by a YAML block when structured information is needed, such as the recipient and content of a message. An agent can issue multiple actions in a single response, allowing it to use each response efficiently.
The dialogue is therefore composed mainly of alternating Station observations and agent responses. When it approaches a configured context limit, generally around 300,000 tokens in this study, the Station asks the agent to write a compact summary of its activities. This summary, together with key messages, is carried into a refreshed context so that the agent can continue its work.
A.2 Agents
Agent composition.
Unless otherwise specified, a Station begins with six agents: two powered by GPT-5.5, two by Claude Opus 4.8, and two by Gemini 3.1 Pro. When an agent leaves, the Station spawns a new agent powered by the same model, keeping the six-agent composition throughout the run.
Lineage.
Agents are organized into lineages. A lineage is a sequence of agents that share a name, private notes, and a continuing research identity. A new agent can inherit an existing lineage of the same model and become its next generation, or create and name a new lineage to begin a different research style. For example, an agent that inherits the lineage of Noesis II becomes Noesis III and gains access to all private notes and records left by Noesis I and Noesis II.
System prompt and role.
All agents receive a shared system prompt describing the Station’s research philosophy, including the standard for a publishable archive paper and the goal of making general scientific contributions. Each agent also receives a specialized research role. Initial roles are sampled from generic templates that each emphasize a different research style: analytical, creative, synthetic, empirical, or strategic. When an agent leaves, it can instead write the role of its own descendant, often giving more task-specific guidance and a more deliberate description of the lineage’s research style. This encourages diverse research behavior across agents while preserving useful differences between lineages.
Agent lifecycle.
An agent can remain in the Station for at most 200 ticks. For its first 40 ticks, it works in isolation, without access to the Station’s communal knowledge or communication with other agents, but with access to the records of its own lineage. This period is intended to encourage independent exploration. The agent then becomes mature and gains access to the main collaborative rooms. At age 100 ticks, it becomes tenured and may choose to leave the Station before reaching its maximum lifetime.
Supervisor.
The Station also appoints a supervisor from time to time. It selects at random a GPT-5.5 agent that has published at least one accepted archive paper. The supervisor provides high-level guidance, encourages agents to explore promising directions deeply, and helps prevent duplication of work, while leaving each agent responsible for its own research. After a supervisor leaves, the Station waits 200 ticks before appointing another supervisor, creating periods of less structured exploration.
A.3 Rooms
The Research Center and the Archive Room are the two main rooms in the Station. Their functions are described below, together with the new Question Room. The remaining rooms are summarized in Table 1.
Research Center.
The Research Center is the Station’s main room for computational experiments. It presents the research task, accepts experiment submissions, runs evaluations, and records their results. It also provides persistent storage for code and artifacts. Agents can review evaluations by their peers and reuse stored code and artifacts, allowing experimental knowledge to accumulate.
To start a Station on a new problem, the user generally provides two components: a task specification and an evaluator. The task specification describes the research problem, submission format, constraints, and evaluation rule. The evaluator is a function that computes a score from an input construction. For example, the kissing-number evaluator takes a proposed set of vectors and reports the total overlap among the corresponding spheres, with zero indicating a valid configuration. Both the task specification and evaluator are available for agents to read.
Agents can also use the Research Center as a sandbox for general computational work. An experiment need not return a construction in the format required by the evaluator; agents can use it for diagnostic calculations, testing conjectures, analyzing earlier results, etc.
Station v2 introduces a separate coder, powered by GPT-5.5 through Codex [59], to help agents implement their experiments. Instead of writing and debugging code itself, an agent submits specific natural-language instructions for one experiment. The coder implements those instructions, runs the evaluator, fixes implementation errors, and returns a report. This allows agents to focus on scientific work, such as designing experiments and interpreting their results, rather than low-level coding work such as debugging.
Archive Room.
The Archive Room is the main knowledge hub of the Station. Agents can publish their findings as archive papers and read papers published by earlier agents. These papers remain available throughout the run, allowing results, methods, and useful negative findings to be passed between agents and accumulated over time. The archive therefore grows throughout the run, gradually expanding the Station’s knowledge of the problem.
Every submitted paper is assessed by a reviewer powered by GPT-5.5. It judges whether the work is rigorous, novel relative to the existing archive, useful to the research goal, and properly supported and cited. Accepted papers are published in the Archive Room, while rejected papers are returned with comments and suggestions so that the author can revise the work or pursue a different direction.
Station v2 also introduces an Archive Surveyor, powered by GPT-5.5 through Codex. As the Archive Room grows to contain dozens or even hundreds of papers, reading the entire literature becomes time-consuming. An agent can instead ask the Archive Surveyor for a literature survey on a particular question or research direction. The surveyor searches the accumulated archive papers and returns a concise survey with citations to the original records. Agents can still read any archive paper directly when they need its full details.
Question Room.
Station v2 introduces a Question Room, where agents can post new research questions and vote on solutions proposed by their peers. The room encourages scientific exploration beyond the main task; for example, solving a related or reduced problem may provide insight into the original problem. Only tenured agents can enter, limiting the time that agents spend away from the main task early in their lifecycle.
Other rooms.
Most other rooms support different forms of communication or reflection. Their functions are self-explanatory and are not described in detail here.
A.4 Mechanisms
Holiday.
Every ninth and tenth tick are declared a holiday. During these ticks, agents cannot run experiments or submit archive papers. Instead, each agent receives a random prompt from a large pool. These prompts encourage broader reflection, such as using metaphors, examining an unexpected observation, revisiting an abandoned idea, or drawing on another field. Most are adapted from the night-science practices described by Yanai and Lercher [38]. The holiday creates regular pauses from routine work in which agents can reconsider their assumptions and explore less obvious directions.
Meta-reflection.
Station v2 also introduces compulsory meta-reflection for mature agents. At least once every 25 ticks, an agent enters the Reflection Chamber and receives a randomly selected high-level reflection prompt. The prompt typically asks GPT-5.5 to act as an external human expert and review the agent’s recent research journey from a different perspective. During this reflection, GPT-5.5 temporarily replaces the agent’s usual model, as we found that it produced the highest-quality reviews. The motivation is to align agents with the broader interests of human researchers, including curiosity, understanding, and scientific value beyond immediate improvement of the evaluation score.
Stagnation protocol.
When the evaluation frontier has not improved for 320 ticks, the Station activates the stagnation protocol. The protocol sends a system message to every mature agent. It randomly assigns each agent one of several lanes: exploration, exploitation, revival, understanding, or strategy. Each lane asks the agent to review the available evidence, question its current assumptions, and develop a different response to the stagnation. The use of multiple lanes encourages diverse paths for escaping scientific stagnation.
Multistart.
Station v2 introduces multistart, which runs eight independent Station rollouts for 40 ticks from the same starting state. A GPT-5.5-powered administrator then compares their progress and selects the branch with the greatest scientific value to continue. Multistart is designed to capture the substantial variation in research trajectories across rollouts. It is used where this variation is expected to be largest: during the first 40 ticks of a Station and the first 40 ticks following activation of the stagnation protocol. The branches are run in parallel, so multistart generally does not increase wall-clock time when sufficient compute is available.
Appendix B Sources of the pre-AlphaEvolve literature column
The pre-AlphaEvolve literature curve of Figure 1 is a reproducible reference assembled from work predating AlphaEvolve. No single paper tabulates these finite values. We therefore take the minimum over the explicitly defined families below, each evaluated at the pair in question.
Bukh–Chao [13], Proposition 11. We use the quadratic-residue block Gn=⋃a{(t,a12+ta1,…,an−12+tan−1):t∈𝔽q} and the recursion Kn=Gn∪(Kn−1+xn), with Kn−1 embedded in a horizontal hyperplane. Proposition 11 makes this Kakeya for every full translation xn. We retain the smallest certified placement found from complete transverse shift histories and from translated horizontal slices, materialize each selected set, and check a complete witness line in every projective direction. This is essential: retaining only one locally best child, or fixing the containing slice, gives larger values at some benchmark pairs.
Mockenhaupt–Tao [57], in the form recorded by Saraf and Sudan [69]. The displayed union has exact size (q−1)(q+12)d−1+qd−1; the sum q(q+12)d−1+qd−1 is a convenient upper bound before subtracting the intersection.
Kopparty, Lev, Saraf and Sudan [42]. Lemma 17 gives the upper bound q∑j<d(q+12)j (its displayed strata may overlap), and the missing-digit construction of Theorem 7 has exact size (q−1)d+2d−1, which is the classical 2d+1−1 at q=3.
Blokhuis–Mazzocca [11]. In d=2 the problem is settled. The minimum is exactly p(p+1)/2+(p−1)/2 for odd p, with a matching construction.
Products. A product of Kakeya sets is Kakeya of exactly the product size, so every product of best bounds in complementary lower dimensions is admissible, with the sharp planar value above as the d=2 factor.
Lev [46]. At p=3, the exact value k3=13 and the bound k4≤27, both from a computer search.
The Bukh–Chao recursion supplies the selected value at all 22 pairs with p≥5. At p=3, the values 13 and 27 are smaller in dimensions 3 and 4, while in dimension 5 the recursive value, the missing-digit construction and 2d+1−1 all give 63. Products never attain the minimum on their own at any pair in range. Comparing the two reference curves against each other, AlphaEvolve is below the pre-AlphaEvolve literature at 18 pairs and the pre-AlphaEvolve literature is below AlphaEvolve at 5, namely both p=3 pairs in d=3,4, and the three larger primes in d=5.
Table 3 reports all 25 benchmark pairs. The initial evaluation is our first evaluation of the pre-AlphaEvolve constructions. The final pre-AlphaEvolve literature column takes the minimum over the families described above after incorporating the extended placement search within the Bukh–Chao recursion. This search improves the initial evaluation at twelve pairs and leaves it unchanged at the other thirteen.
| (d,p) | Initial Evaluation | Pre-AlphaEvolve Literature | AlphaEvolve | Station |
|---|---|---|---|---|
| (3,3) | 13 | 13 | 15 | 13 |
| (3,5) | 53 | 53 | 53 | 53 |
| (3,7) | 129 | 129 | 128 | 128 |
| (3,11) | 440 | 440 | 438 | 437 |
| (3,13) | 699 | 698 | 697 | 697 |
| (3,19) | 2,034 | 2,034 | 2,031 | 2,030 |
| (3,23) | 3,509 | 3,509 | 3,505 | 3,504 |
| (3,29) | 6,837 | 6,837 | 6,833 | 6,833 |
| (3,31) | 8,295 | 8,295 | 8,290 | 8,288 |
| (3,37) | 13,867 | 13,866 | 13,861 | 13,861 |
| (3,41) | 18,709 | 18,708 | 18,701 | 18,701 |
| (3,43) | 21,504 | 21,504 | 21,495 | 21,495 |
| (3,47) | 27,899 | 27,899 | 27,892 | 27,889 |
| (3,53) | 39,687 | 39,686 | 39,677 | 39,677 |
| (4,3) | 27 | 27 | 31 | 27 |
| (4,5) | 164 | 163 | 162 | 161 |
| (4,7) | 529 | 528 | 527 | 527 |
| (4,11) | 2,689 | 2,689 | 2,687 | 2,684 |
| (4,13) | 4,973 | 4,972 | 4,966 | 4,962 |
| (4,17) | 13,524 | 13,521 | 13,514 | 13,509 |
| (4,19) | 20,593 | 20,586 | 20,583 | 20,579 |
| (5,3) | 63 | 63 | 63 | 53 |
| (5,5) | 503 | 497 | 510 | 490 |
| (5,7) | 2,145 | 2,142 | 2,187 | 2,135 |
| (5,11) | 16,348 | 16,307 | 16,427 | 16,288 |
These twelve changes do not alter the comparison tally: the Station remains strictly smaller than the better reference at 14 pairs, tied at 11 and worse at none.
The construction of every candidate above, and the check that each is Kakeya, are carried out in the accompanying notebook.
来源:Hacker News 热门(buzzing.cc 中文翻译) · arxiv.org