The U.S. Department of Energy published a short bulletin on August 7 asking universities, companies, national labs and research nonprofits to hand over scientific data, trained models and evaluation material for a new family of open-weight AI models. The first application window closes on August 14.
That leaves five days from today for an institution to decide whether it wants its research corpus inside a federal training program, which is not much runway for anything that has to pass a legal review first.
The program is the Genesis Open Models Initiative, hosted by Argonne National Laboratory. Its first model is Genesis-Science-1, or GS1, built with the San Francisco lab Arcee AI.
The announcement isn’t the model. It’s the recruitment drive.

GS1 itself was announced back on July 22, in a joint release from DOE and Arcee. What happened on August 7 is different, and the distinction matters if you’re deciding whether to apply.
The August bulletin came from the Office of the Under Secretary for Science, and it widens the program past a single company. DOE says it wants open-weight models from other organizations too, both as base models for downstream fine-tuning and for immediate deployment, with what it calls transparent provenance.
It also wants curated scientific datasets and specialized corpora for future pretraining cycles, and teams willing to build domain-adapted versions of GS1 for specific fields.
So Arcee is the first industry partner, not the last one. DOE’s portal states plainly that it’s still accepting new partners for the open models effort.
There are two tracks for the 2026 program. Foundation-stage data (scientific text, code, documentation, structured technical collections) is due August 14, with delivery by August 28 if selected.
Post-training material (fine-tuning examples, workflow environments, reinforcement-learning tasks, held-out evaluations, rubrics and verifiers) is due August 25, with delivery by September 14. Additional windows are expected every three months.
One thing to check before you apply: Arcee’s own announcement page still lists the earlier deadlines of August 6 and August 20, which have shifted. DOE’s portal and the energy.gov bulletin both say August 14. Go with the DOE dates.
Applications don’t transfer any actual scientific material. The public form collects descriptions and metadata only, and each contributor states its own proposed terms of use.
Submissions then pass through five review gates: scientific fit, rights and handling, expert and evaluation readiness, technical integration, and final selection. Selected contributors get early evaluation access as the system develops, plus credit in the technical report.
Why a federal agency is recruiting open-weight models at all

This is the part the bulletin doesn’t spell out, and it’s the reason the timing makes sense.
America’s position in open-weight AI has slipped badly. The ATOM Project, which tracks open-model activity, found that from late 2023 through March 2026 roughly 70% of newly created derivative open models worldwide were built on Alibaba’s Qwen.
Llama, which held about 40% two years earlier, had dropped to around 10%. AI researcher Nathan Lambert wrote in that report that the U.S. has already fallen behind on both performance and adoption.
Independent tracking of OpenRouter traffic put Chinese open-weight models at roughly 61% of all tokens consumed by May 2026.
Now add the procurement side. DeepSeek is restricted on many U.S. government devices, and states including Texas, New York and Virginia banned it on state hardware starting in early 2025. A national lab that wants to run a capable open model on its own hardware, inside its own security boundary, is choosing from a short and politically constrained menu.
That’s the gap DOE is trying to fill. Not a chatbot. An American model that a lab can hold, version, freeze for years, and adapt without a permanent dependency on someone else’s API.
Darío Gil, DOE Under Secretary for Science and the Genesis Mission’s director, framed GS1 as “an open model trained on the real work of the national laboratories” and answerable to the scientists doing that work.
Arcee CEO Mark McQuade put the strategic version more bluntly in the July announcement, arguing that a country can’t lead in AI if everything it leads in is closed.
The parent program launched by executive order on November 24, 2025. The Genesis Mission connects DOE’s 17 national laboratories, its leadership-class computing facilities and its data resources, with a stated goal of doubling the productivity and impact of American science and engineering within a decade. DOE announced $320 million in investments to the labs to advance it. [Link to prior coverage of the Genesis Mission executive order.]
What GS1 is actually being trained to do
The technical description is more interesting than the usual agency press language, mostly because it describes a working environment rather than a benchmark.
Scientific computing, as DOE describes it, rarely starts with a clean prompt. A researcher inherits an aging Fortran codebase, a half-finished simulation campaign, contradictory run logs, and several defensible options for what to try next.
GS1 will train inside workbenches built to reproduce exactly that mess, covering high-performance-computing code modernization, experimental analysis, simulation campaigns, materials science and energy systems. Training environments may include Python, Fortran, C and C++, MPI and OpenMP, CUDA and HIP, notebooks, simulation packages and job schedulers.
The model runs through a governed execution system. Tools execute in sandboxed, staged environments, and the system records prompts, tool calls, code changes, datasets, intermediate artifacts and conclusions for each run. Humans approve anything touching safety, security, publication or resource use. DOE states that GS1 will not receive blanket access to its systems.
The success criterion is unusual, and it’s the most interesting design decision here. A run counts as successful only if it carries a workflow from plan through report, revises when the evidence changes, recovers from tool failures, and leaves a record another researcher can inspect.
Scientists judge whether the result is sound and whether the evidence supports reproduction. That’s a reproducibility standard, not a leaderboard score.
Arcee’s track record is why DOE picked a 30-person company over a hyperscaler. The lab has raised just under $50 million total, and in early 2026 it spent roughly $20 million of that, close to half its capital, on a single 33-day training run across 2,048 NVIDIA B300 GPUs
. The result was Trinity Large, a sparse mixture-of-experts model with about 400 billion total parameters and roughly 13 billion active per token, released under a permissive license. [Link to prior coverage of Arcee’s Trinity Large release.]
The questions the bulletin leaves open
There’s no money on the table. That was the first thing commenters on Hacker News noticed, where the announcement drew 345 points and nearly 150 comments. Contributors are asked to supply curated data, expert reviewer time and evaluation environments in exchange for early access and a credit line in the technical report
. One commenter, who appeared to be writing from inside a lab environment, suggested the obvious fix: attach funding for a postdoc or a graduate student to each accepted contribution, and teams would compete to get in.
There are also no published specs, and the one number in circulation didn’t come from the program documentation. Arcee CTO Lucas Atkins wrote on X on July 22 that GS1 is being developed as a trillion-parameter-class model for high-difficulty scientific workflows, a line Arcee has repeated on social but hasn’t put on either the DOE portal or its own announcement page.
Alan Thompson’s Memo, which tracks large model releases, puts delivery later this year. Arcee’s official page says only that its operating experience lets it deliver GS1 on an accelerated schedule.
So: no confirmed parameter count, no architecture detail, no training data disclosure, no firm release date on any DOE page. At this stage GS1 is a program and a call for participation, not something you can download.
If you’re weighing a contribution, you’re being asked to commit data to a model whose scale you know about from a social post.
And “foundation model” doesn’t necessarily mean a language model. As one commenter pointed out, neither the DOE bulletin nor the GS1 page uses the words “LLM” or “language” anywhere, and a fair number of proposals under the wider Genesis Mission answer the foundation-model call with non-language architectures.
The GS1 description does read like a language model wrapped in an agentic harness. But nobody has said so outright.
The sharpest criticism is about relevance. An open model matters only if researchers actually choose it, and American open releases have struggled to hold attention for long. Nvidia’s Nemotron series has been briefly competitive at each launch without ever leading, and Ai2’s fully transparent Olmo work has more admirers than users.
A government model that ships to polite indifference does nothing for the problem it was built to solve.
The early signs are quiet. Two days after the bulletin, X activity on the initiative is mostly link-sharing from news bots and aggregator accounts, with Arcee’s own post the highest-performing item at 68 likes. No thread has broken into real discussion. For a program asking institutions to commit data inside a week, that’s a thin start.
What to watch
August 14 is the immediate one, followed by an August 28 delivery deadline for anything selected in the foundation-stage track. The post-training window closes August 25. After that, the schedule turns quarterly.
Two signals will tell you whether this is real. The first is whether DOE names additional industry partners in the next couple of cycles, since the portal is explicitly open to them and a one-vendor program is a fragile one. The second is what the eventual technical report discloses about training data, because a federally backed model that documents its corpus honestly would be doing something no frontier lab has been willing to do.

