8 minute read

TL;DR: On August 31 the Department of War put OpenAI’s ChatGPT Mil and Starshield AI’s Grok for Government on GenAI.mil, alongside Google Gemini. Three frontier models, one accredited platform, 1.7 million users onboarded out of roughly 3 million people. That is a genuinely fast enterprise rollout and the team that did it deserves credit. But adoption is the easy number. The department still has no fielded, common way to measure whether any of those three models is actually good at the work people are doing with them. The evaluation harness that would answer that question was a solicitation in March. We are scaling access faster than we are scaling judgment, and I have seen where that bill comes due.
What actually happened this week
On Monday, August 31, 2026, the Department of War announced that ChatGPT Mil and Grok for Government are live on GenAI.mil. Google Gemini has been on the platform since December 2025. DefenseScoop reported that all three cleared Impact Level 5, the authorization tier for non public, sensitive unclassified data. TechCrunch put the platform at 1.7 million unique users onboarded against a population of about 3 million military, civilian, and contractor personnel. DefenseScoop had reported 1.5 million active users in June.
ChatGPT Mil is pitched at document heavy unclassified work: planning, policy, logistics, administration. Grok for Government brings reasoning modes and reusable playbooks. Both are unclassified enterprise tools, not the classified capabilities being developed separately for combat operations. That distinction gets lost in the coverage.
Nine months from launch to more than half the department. I have watched enough enterprise IT rollouts die quietly at the pilot stage to say plainly that this one moved. The hard part of adoption, getting the accreditation and the account provisioning and the terms of use to line up, was done well.
The department framed the expansion around choice. The official statement was that it “will continue to build an architecture that prevents AI vendor lock and ensures long-term flexibility for the Joint Force.” I agree with that goal completely. I want to talk about what has to be true underneath it for the goal to mean anything.
Adoption is the easy number
Every program I have supported eventually produces a slide with a user count on it. User counts are cheap to collect and they feel like progress. They are not capability.
1.7 million accounts tells you people logged in. It does not tell you:
- What tasks they actually brought to the model.
- Whether the output was correct.
- Whether the user caught it when it was not.
- Whether a different one of the three models would have done better on that same task.
- Whether any of it changed a decision.
That last one is the only question that matters operationally, and it is the hardest to instrument. If you cannot answer the first four, you have no path to the fifth. Right now the honest answer across the enterprise is that we are largely inferring quality from the fact that people keep coming back.
Return usage is a real signal. It is not a measurement. The gap between user satisfaction and mission effectiveness has embarrassed more than one program at the operational test event.
The test infrastructure was built for a different problem
This is the part I think deserves more attention than it gets.
CDAO has real test and evaluation machinery for AI. The published T&E frameworks cover model T&E, human systems integration T&E, systems integration T&E, and operational T&E. The Joint AI Test Infrastructure Capability, JATIC, builds actual tooling rather than policy PDFs, which I appreciate. The Explainable AI Toolkit and the Natural Robustness Toolkit are the two I point people to when they ask where to start.
Then read the JATIC program documentation and note the current focus: testing AI models for computer vision classification and object detection, with other modalities such as natural language processing and generative AI described as work for future program years.
That is not a knock on JATIC. Computer vision was the right first target, and that work is genuinely hard and valuable. It is a statement about timing. The department’s most mature AI test tooling was built for the AI problem of five years ago, while 1.7 million people are now typing prompts into large language models on the enterprise platform.
Capability got ahead of the test infrastructure. That happens. What you do about it in the next twelve months is the whole ballgame.
The yardstick is still a solicitation
The department knows this, and to its credit it wrote the requirement down. On March 11, 2026, the Defense Innovation Unit and the Office of the Director of National Intelligence released a solicitation called MYSTIC DEPOT, seeking an evaluation harness and government specific benchmarks that together enable “rigorous, reproducible, vendor-agnostic assessment of any AI system against government-defined criteria.” Solution briefs were due March 24 under a Commercial Solutions Opening.
Whoever wrote that requirement understands the problem. It asks for a model interface and execution engine, measurement and scoring, automated red teaming with adversarial prompts, simulation of operational stress and degraded networks, human machine teaming interfaces so subject matter experts can review results, and benchmarks built to resist gaming. It asks for all of it across unclassified, secret, and top secret environments. That is close to the list I would have written.
I could not find a public award announcement. So as of this week, the sequence reads: three frontier models fielded to more than half the department, and the vendor agnostic yardstick for comparing them still working its way through acquisition.
That ordering is not unusual in defense. It is also precisely the ordering that produces the accreditation fights, the trust deficits, and the awkward operational test findings that show up two years later, long after the launch press release.
Vendor choice is not the same as vendor comparison
Putting three vendors on one platform is a procurement achievement. It is not yet a competitive one.
Competition requires that you can tell which one is better at something specific. Without a common benchmark suite tied to real mission tasks, three models in a portal is three menu items, not a competition. Nobody can say which model to use for a logistics rollup versus a policy comparison versus a contract summarization, so users pick by habit, by interface preference, or by whichever brand name they recognize. That is how you get de facto lock in on a platform explicitly designed to prevent it.
CDAO awarded contracts worth up to $200 million each to Anthropic, Google, OpenAI, and xAI in July 2025 to develop agentic AI workflows. That was a bet on breadth. Breadth without measurement is just a bigger surface area to defend.
The pattern is bigger than one platform
In April 2026 GAO published report GAO-26-107859, “Artificial Intelligence Acquisitions: Agencies Should Collect and Apply Lessons Learned to Improve Future Procurements.” It reviewed 13 AI acquisitions across the Department of Defense, DHS, GSA, and VA. The finding that stuck with me: none of the four agencies had policies requiring officials to systematically collect lessons learned from AI acquisitions. GAO also found a recurring problem getting AI technical experts in the room to evaluate proposals, and unclear cost structures. All four agencies concurred with the recommendations.
Read that next to an earlier GAO finding that generative AI use cases across federal agencies grew roughly ninefold in a single year, from 32 to 282, and the shape of the problem is clear enough. We are getting very good at acquiring and deploying AI. We have not built the institutional memory or the measurement discipline to tell which of it worked.
So What
If you are running a program, a portfolio, or a company that sells into this, here is what I would take from the week:
- Stop reporting adoption as if it were performance. Put a user count on the slide if you must, but never without a task level accuracy or quality measure next to it. If you do not have one, say so out loud. That is a finding, not a gap to paper over.
- Write the evaluation plan before the pilot, not after. Every fielded model needs a test plan, not a demo. Define the mission tasks, the scoring criteria, and who adjudicates a disagreement, and do it before anyone gets an account.
- Build your own small benchmark now. You do not need MYSTIC DEPOT to be awarded to assemble 50 real tasks from your own workflow with known good answers and run all three models against them quarterly. That is a week of work and it will outperform any vendor claim you are handed.
- Instrument the human, not just the model. The failure mode in document heavy staff work is not a hallucinating model. It is a rushed reviewer who accepts a fluent wrong answer. Measure catch rates.
- If you sell here, bring test evidence. The vendors who show up with reproducible evaluation results against government relevant tasks are going to have a structural advantage the moment a real harness exists. Start generating that evidence now.
None of this is an argument for slowing down. The rollout was the right call and the speed was impressive. It is an argument that the measurement work is now the long pole, and the department has already told us what it wants that work to look like. The task is to go build it.
Three models on the platform is a good headline. The better one, whenever we earn it, will be the day someone can say which model to use for which job, and show the data behind the answer.

