The Pentagon Is Counting AI Users. That Is Not the Same as Counting Value.

Navy-and-gold illustration showing defense AI activity passing through verification gates toward a trusted mission outcome.

8 minute read

TL;DR: More than 1.6 million people used GenAI.mil in its first six months, generating tens of millions of prompts and hundreds of thousands of agents. The Department’s performance plan separately targets a 75 percent monthly-active-user adoption rate and a cumulative 15 percent productivity gain in designated AI pilots in FY27. Those are useful measures of reach and speed. They are not yet a value model. Time saved becomes defensible cost avoidance only after quality, rework, failed attempts, released capacity, total cost, and uncertainty are counted against a traceable baseline. The next phase of defense AI needs an evidence ledger, not a larger activity counter.

This post uses public sources and a framework drawn from general professional experience. It contains no nonpublic program, platform, customer, or employer data.

The Department has cleared a large-scale adoption hurdle: people are willing to try the tools.

That matters. Defense organizations have buried plenty of software beneath low adoption and mandates nobody followed. GenAI.mil cleared that bar quickly.

Now comes the harder question every executive, program manager, and taxpayer should ask: what did the usage produce?

The Adoption Numbers Are Real

When CDAO began transitioning the GAMECHANGER policy corpus into GenAI.mil in June, it reported that more than 1.6 million Department personnel had used the platform, generating tens of millions of prompts and deploying hundreds of thousands of agents in six months. That is cumulative reach, not the same thing as sustained active use.

The latest Department performance plan turns that momentum into formal targets. It defines adoption as monthly active GenAI.mil users divided by civilian and military personnel with network access, with targets of 40 percent in FY26 and 75 percent in FY27. Separately, it targets an average workflow-productivity improvement within designated AI pilot projects of 5 percent in FY26 and 15 percent cumulative in FY27, measured by reduction in task or process time.

Those are reasonable first measures: adoption shows reach, active use shows return, and cycle time shows speed. Each stops one step before value.

A login is not an outcome. A prompt is not a product. An agent deployment is not a capability. And a task completed faster is not automatically a dollar saved.

As I argued in AI at Wartime Speed, generative AI has crossed from pilot to infrastructure. That makes a defensible value measure more important, not less.

Time Saved Is Capacity, Not Cash

This distinction matters because AI business cases routinely turn minutes into money too quickly.

If an analyst cuts a four-hour task to two, the workflow releases two hours of capacity. That is real, but payroll, contract expenditure, or another cost base has not necessarily fallen.

The released time becomes one of three things:

  1. Cost savings if spending actually falls—for example, less overtime, fewer purchased labor hours, or a reduced contract requirement.
  2. Cost avoidance if the organization absorbs new demand without adding people or money it otherwise would have needed.
  3. Capacity gain if the same workforce completes more mission work, improves quality, or tackles a backlog.

All three are valuable. They are not interchangeable.

When I built an enterprise AI value model, the important decision was to keep those categories separate and trace every claim to a use case, baseline, attributed hours, labor rate, and accepted output. The number is smaller than multiplying prompts by guessed savings. It is also one I would defend in a budget review.

Quality Is the Missing Denominator

Speed without quality is how an AI pilot creates invisible rework.

A draft may arrive in seconds and still consume an hour of fact-checking. A polished summary may omit the exception that controls the decision. If measurement stops when the model answers, the dashboard credits the tool for work a human still has to redo.

NIST is working one critical part of this problem. Its agentic-AI research uses evaluation probes and machine-readable audit trails to test whether claims are faithful, complete, and sufficiently supported. The useful pattern is that evaluation travels with the output instead of disappearing after a pilot.

The basic unit of AI value should therefore be an accepted, fit-for-purpose output, not a generated output.

Once that denominator is in place, the useful measures change:

  • first-pass acceptance rate;
  • reviewer time per accepted product;
  • material error and omission rate;
  • rework hours;
  • provenance and citation completeness;
  • regression after a model, tool, prompt, or source corpus changes; and
  • drift over time or as the input distribution changes.

Now the productivity number means something.

A Five-Line Evidence Ledger

The Department needs a small, consistent record for high-value workflows. For each use case, capture five lines:

1. Baseline

Define the unit of work, cycle time, quality, volume, backlog, and labor mix before AI. If the baseline moves, report the range and uncertainty.

2. Attributable Change

Compare representative matched work units with and without the tool. Where practical, randomize or phase rollout across similar teams so task mix and early adopters do not decide the result. State the sample, observation window, and uncertainty. Separate user, reviewer, rework, and failed-attempt time.

3. Acceptance and Rework

Record whether each output was accepted, corrected, or rejected and count the labor needed to make it usable. A two-hour draft that creates three hours of remediation is not a gain.

4. Realized Value

State what happened to released capacity: reduced purchased hours, avoided hiring, cleared backlog, greater throughput, a shorter intelligence cycle, or a better decision. Do not claim savings when the evidence supports capacity or avoidance.

5. Total Cost and Risk

Include licenses, compute, integration, training, governance, review, remediation, and maintenance of approved data and agents. Record security, privacy, and mission risks beside the financial result.

A practical financial view starts by comparing the same volume of acceptable work:

Realized capacity benefit = net verified hours avoided across all attempts × loaded labor rate × realization factor

Net verified hours include all user, reviewer, rework, rejected-output, and abandoned-attempt labor needed to produce the baseline volume of acceptable work. The realization factor is the share of released time that became productive mission capacity. Keep it conservative and visible.

Then calculate the business case in ordinary terms:

Net value = realized benefit − total allocated cost

ROI = net value ÷ total allocated cost

Allocate costs to the same use case and period. A hypothetical team of 50 analysts verifying two net hours avoided each week for 46 weeks at an $85 loaded rate creates $391,000 in gross capacity. At an assumed 70 percent realization factor and $90,000 allocated cost, illustrative net value is $183,700 and ROI is about 204 percent. It becomes defensible only when the rate, period, realization, costs, quality threshold, and uncertainty are evidenced—preferably as ranges. Call it savings only if spending fell.

Mission value belongs beside that equation, not forced inside it. A faster targeting decision, reduced unscheduled maintenance or repair time, or a better-supported command option may matter far more than the labor hours involved. Report the mission measure directly.

Agents Raise the Stakes

The distinction between activity and value becomes more important as the Department moves from chat to action.

GenAI.mil users have already deployed hundreds of thousands of agents. Separately, the Department’s Agent Network is designed to scan intelligence and operational systems and present targeting options to commanders within seconds. The Department says humans retain the decisions and the system will receive rigorous testing and operational evaluation.

Good. Because an agent does not merely generate more text. It can select tools, query data, take intermediate actions, and compound error across steps.

The metric cannot be “agent ran successfully.” It needs to show:

  • which data and tools the agent used;
  • whether it stayed within its authority and resource limits;
  • where a human intervened;
  • whether the evidence supported the recommendation;
  • whether the final action improved the mission measure; and
  • how behavior changed after any model, prompt, tool, or data update.

This is ModelOps as an evidence discipline. The model version matters. So do the workflow, sources, tools, evaluator, operating conditions, and decision that followed.

The evidence chain is only as trustworthy as the data beneath it. A renamed platform does not fix ownership, lineage, or authoritative-source disputes, as I wrote in The Pentagon Renamed Its Data Platform. The Data Problem Came With It.

The Honest Counterweight

Measurement can become its own form of failure.

If every five-minute use of AI requires a 20-field form, people will route around the process or stop experimenting. A perfect control system delivered in two years would also miss the reason the Department moved to enterprise AI in the first place.

The answer is proportionality.

Instrument routine activity automatically. Sample low-risk workflows. Require stronger baselines, evaluation, traceability, and human review as consequence and autonomy increase. Reuse approved measures across repeated workflows. Let the evidence burden rise with mission risk.

The Department should keep counting users and cycle time. Those measures are leading indicators. It should stop treating them as the finish line.

So What

  • If you lead an AI portfolio: run two dashboards. One shows active adoption, availability, activity, and reuse of validated assets with revalidation status. The other shows accepted outputs, quality-adjusted hours, realized capacity, mission measures, total cost, and uncertainty.
  • If you manage a program: define the baseline and decision outcome before the pilot. Decide in advance what would count as savings, avoidance, capacity, and failure.
  • If you build agents: make provenance, tool-use logs, versioning, evaluation results, and human interventions part of the product. An audit trail added later will always be incomplete.
  • If you approve budgets: ask for the evidence chain behind the ROI number. A precise dollar figure without a traceable baseline is theater with decimals.

The Department has demonstrated large-scale willingness to try AI. The next competitive advantage is knowing, with evidence, which uses deserve to scale.

Count the users. Fund the value.

Dr. Shane Turner
Dr. Shane Turner

Exploring ideas and challenging assumptions about defense technology, one post at a time.

Get new posts in your inbox

We don’t spam! Read our privacy policy for more info.

Share