Start with the TSA number
Salesforce marked one year of Missionforce on September 16 with a post full of scale, and most of the attention will go to the contract sizes. The U.S. Army awarded a $5.6 billion, 10-year IDIQ. The Air Force 441st VSCOS manages a $13.5 billion fleet of more than 84,000 vehicles across 389 global locations. An Army Human Resources Command deployment covers 9.2 million soldiers, veterans, civilian staff and families. Salesforce says the defense sector business has grown 80% year over year. Those figures tell you who signed, and not much else.
The line that should stop an enterprise architect is smaller and sits further down. Salesforce says TSA's Ace AI agent handles roughly 100,000 traveler conversations monthly, resolves 96% of them without escalation, and reduces interaction costs by more than 90%. A production autonomy rate published against a stated monthly volume is uncommon. Most commercial AI case studies hand you a percentage with no denominator, or a denominator with no percentage.
The useful part of this release has nothing to do with whether you sell to government. It gives you a reference point to hold your own numbers against.
Why government releases make better evidence
Public sector procurement forces disclosure that commercial contracts hide. Salesforce can name TSA and the 441st VSCOS because those programs already exist in the public record, with appropriations and oversight attached. A bank running a comparable deployment shows up in a case study as a large North American financial institution with an undisclosed volume and a satisfaction uplift.
If you are assembling a business case, that difference matters more than vendor preference. The comparable you want for a move from pilot to practice is a named workload with a stated volume at a named agency. Our read is that a number attached to a federal program carries a longer tail of scrutiny than one in a marketing PDF, because somebody can ask about it later in a setting the vendor does not control. The same logic is why we read public sector demand signals from Workday closely.
None of that makes the number true. It makes it checkable, which is more than the market usually offers.
The risk profile behind the 96%
This is the part that needs care. The release says traveler conversations. It does not say what those conversations are about, and that gap decides how much the 96% is worth to you.
An agent answering questions about checkpoint hours or what somebody can carry through screening has a failure mode you can absorb. A wrong answer gets corrected at the podium, and nobody loses an entitlement over it. An agent resolving 96% of disability claim determinations without a human in the path would be a different story, with a different appeals process behind it. Same headline percentage, very different consequence per error.
So when you carry this figure into your own planning, carry the workload class with it. The containment rates coming out of ServiceNow deployments have the same property. A resolution rate only compares across workloads of similar consequence, and vendors rarely publish the consequence.
The missing 4% is the design question
Four percent of roughly 100,000 conversations a month lands somewhere near 4,000 escalations. The release does not describe what happens to them, and that is the part we would ask about first.
Does the traveler reach a staffed queue with a response target, or get handed a phone number, or get dropped onto a page of FAQs? Who owns that queue, and does the escalation carry the conversation history or start cold? Our guess is that the design exists and simply is not in an anniversary post, but the escalation path is where reputational damage collects in every agent deployment we have looked at. A well built handoff between human and agent work decides whether 4,000 people a month get helped or get stuck.
There is a second definitional question. Resolved without escalation can mean the traveler got what they needed, or it can mean the traveler gave up and walked away. Abandonment registers as resolution in a lot of contact center instrumentation. The post does not say which definition TSA uses, and the answer would change how you read that 90% cost reduction too.
The OpenAI piece is an announcement so far
Salesforce also announced a new Missionforce partnership with OpenAI, described as bringing advanced AI models and accelerated computing into secure government environments. That sentence is close to the whole of it. No model list, no accreditation boundary named, no timing.
For anyone running a model gateway, the questions write themselves. Which models, under which authorization, with what inspection at the boundary, and what happens to prompt and response data once it sits inside an accredited environment. Those are the same questions you should already be asking about runtime policies at your own AI gateway, and in government the answers arrive as documentation long after the press release.
Kendall Collins, CEO of Missionforce and Government Cloud, said, "One year in, the momentum behind Missionforce shows how quickly government agencies are moving from AI ambition to operational deployment." Read against the TSA figure, that holds up for one workload class. Whether it holds at the next level of consequence is the open question.
What to ask before you cite this
If somebody brings the 96% into your architecture review as evidence that autonomy is ready, push on a few things before you nod along. Ask what the conversation class is. Ask what counts as resolved. Ask where the escalations go, who staffs that path, and what the response target is.
Then ask the harder one, which is whether the workload you want to automate carries the same consequence per error as a traveler asking about a checkpoint. If it does not, the 96% tells you nothing about your own ceiling. If it does, ask your own vendor why they will not publish a comparable number with a denominator attached.
Bring one figure to your next review. What share of your own agent conversations closed without a human last month, and what share of those closed because the person stopped trying? If nobody in the room can answer the second half, you do not have a resolution rate yet.



