Defense contractors chasing CMMC 2.0 compliance for their AI systems just watched the ground move under them. DoD suspended the transition to CMMC Phase 2 on July 13, 2026, less than four months before it was supposed to start, which means half the "November 2026 deadline" content still circulating is now wrong. Phase 1 is live and enforced. Phase 2 is under review. Neither fact changes what you're required to do the moment an AI workload touches Controlled Unclassified Information.
This guide covers where CMMC actually stands today, why commercial AI tools are a non-starter for CUI regardless of Phase 2's status, and how to build a self-hosted GPU deployment that keeps you in scope. For the FedRAMP mechanics behind why no GPU cloud can legally touch CUI yet, see our FedRAMP GPU cloud buyer's guide, which covers the same 32 CFR Part 170 rule from the civilian-agency side.
What CMMC 2.0 Actually Requires for AI Workloads Touching CUI
If an AI system processes, stores, or transmits CUI, it's in scope for CMMC, full stop, regardless of what phase the program is in. The AI part doesn't create a separate rulebook. It has to sit inside the same access control, encryption, logging, and incident-response boundary as every other system that touches controlled data (NR Labs, "AI and CMMC: What Defense Contractors Need to Know Now").
Phase 1 Is Live, Phase 2 Just Got Suspended (July 13, 2026)
CMMC Phase 1 began November 10, 2025. Every new DoD solicitation now requires contractors to complete an annual Level 1 self-assessment against the 15 basic safeguarding requirements in FAR 52.204-21 and submit an affirmation to the Supplier Performance Risk System. Contractors can't win new business without that affirmation on file (Dorsey & Whitney, CMMC client alert, November 2025). Under the original schedule, Phase 1 runs through November 9, 2026.
Phase 2 was supposed to layer on top of that: C3PAO-certified Level 2 assessments for contracts involving CUI, starting November 10, 2026. On July 13, 2026, DoD halted that transition and launched a 60-day CMMC Reform Task Force review, with a public RFI comment window open through August 14, 2026 (DefenseScoop, "DoD Halts CMMC Cybersecurity Requirements Phase 2," July 13, 2026). DoD CIO Kirsten Davies laid out the reasoning bluntly: "the math just simply doesn't math for small to medium-sized businesses to even get compliant by the transition date." DoD's own estimate put full Level 2 compliance at more than $7 billion a year for small and medium defense businesses, a cost the department cited as a core reason for the pause. Under Secretary of Defense for Acquisition and Sustainment Michael Duffey framed the risk in market terms, warning the original timeline was "forcing companies out of the market at a time when we need them most."
That reasoning tracks with the readiness data. As of March 2026, only about 1,074 organizations nationwide had achieved CMMC Level 2 certification, roughly 1.3% of the estimated 80,000 Defense Industrial Base contractors expected to eventually need it (Secureframe, "CMMC Ecosystem by the Numbers"). The assessor pipeline was never going to clear that backlog on the original schedule either: just over 100 C3PAOs are authorized to conduct Level 2 assessments as of mid-2026, against a Defense Industrial Base that runs well into six figures (Secureframe, C3PAO directory). Redspin's Thomas Graham put the enforcement lag in plain terms before the pause: "I expect to see more and more organizations waking up," describing companies that assumed they had until late 2026 to start taking it seriously (DefenseScoop, "Pentagon Begins Enforcing CMMC Compliance, but Readiness Gaps Remain," November 10, 2025).
What this means in practice: during the suspension, DoD continues to enforce NIST SP 800-171 Rev 2 through self-assessments and select government-led assessments, and Phase 1 requirements stay in effect unchanged (DefenseScoop). If your contract already requires Level 2 self-attestation, that obligation doesn't disappear because third-party certification is paused. Treat the suspension as a reprieve on the assessor bottleneck, not a reprieve on the underlying security requirements.
The 110 NIST 800-171 Controls That Govern Any System Touching CUI
CMMC Level 2 isn't a new standard. It's third-party verification bolted onto a requirement contractors have technically carried since 2017, when DFARS 252.204-7012 took effect. Full Level 2 compliance means implementing all 110 security requirements across the 14 control families in NIST SP 800-171: access control, audit and accountability, awareness and training, configuration management, identification and authentication, incident response, maintenance, media protection, personnel security, physical protection, risk assessment, security assessment, system and communications protection, and system and information integrity.
An AI system doesn't get a carve-out from any of those families. If it processes CUI, it needs documented access control (who can query it), audit logging (what was asked and returned), configuration management (what model version, what patches), and incident response procedures, the same as a file server or a database.
Why an AI Tool Is a Cloud Service Provider the Moment It Touches CUI (32 CFR Part 170)
Under 32 CFR Part 170, any cloud service that processes, stores, or transmits CUI is treated as a Cloud Service Provider, and CSPs handling CUI must carry FedRAMP Moderate authorization or an equivalent (Cloud Security Alliance, "Securing AI in CMMC Level 2 Environments," January 2026). That rule doesn't care whether the "cloud service" in question is a database, a file share, or an LLM API. The moment CUI reaches it, it's an in-scope CSP under the same regulatory test as everything else your SSP already accounts for.
This is the fact that collapses most of the "just use our AI platform, we're compliant" sales pitches contractors hear. As of mid-2026, no neocloud or GPU-specific infrastructure-as-a-service provider holds an active FedRAMP ATO. The only FedRAMP High GPU capacity that exists today sits inside AWS GovCloud, Azure Government, Oracle Government Cloud, and Google Cloud Assured Workloads, each with its own product-scope caveats. A commercial GPU rental platform without that authorization, however good its SOC 2 report looks, is not a legal destination for CUI.
Why Commercial AI APIs Are a Non-Starter for Controlled Data
The fastest way to fail a CMMC assessment isn't a missing firewall rule. It's an employee pasting a paragraph of a CUI document into ChatGPT to save ten minutes on a summary.
The Spillage Problem: Pasting CUI Into ChatGPT or Copilot Is a Reportable Incident
Once CUI enters a commercial AI platform, that text can be stored on infrastructure the organization doesn't control, reviewed by people outside the organization, and potentially retained in ways the contractor has no visibility into (StealthTech365, "Can You Use Copilot or ChatGPT Without Leaking CUI?"). Standard commercial ChatGPT, Gemini, and consumer Copilot don't carry the FedRAMP Moderate authorization DFARS 252.204-7012 requires for CUI-handling systems, so uploading CUI into any of them is a direct compliance violation, not a gray area.
The part that makes this genuinely dangerous is the detection gap. A misconfigured firewall or an unpatched server shows up in a scan. A pasted paragraph doesn't. There's no network signature, no alert, nothing for a SOC to catch. The data leaves the moment someone hits enter, and it's a reportable incident regardless of whether it was caught. If a document falls inside the CUI boundary, it stays in the contractor's enclave, not in a public prompt box, no exceptions for convenience.
What DFARS 252.204-7012 Actually Obligates You to Do
DFARS 252.204-7012 has required "adequate security" for covered defense information since 2017, defined as implementing NIST SP 800-171 and reporting cyber incidents within 72 hours of discovery. CMMC formalizes verification of that same obligation; it doesn't create a new one. CMMC doesn't ban AI outright, but any AI tool that touches CUI has to operate inside the same access control, encryption, logging, monitoring, and incident-response controls as every other in-scope system (NR Labs). CSA's guidance for AI inside CMMC Level 2 environments is specific about what that looks like in practice: network segmentation isolating AI workloads, no unvetted external connectivity for self-hosted AI, encryption in transit and at rest, disabled telemetry, full prompt-and-response logging, and RBAC plus MFA restricted to trained, cleared personnel (Cloud Security Alliance). That's the actual checklist. Everything below in this guide builds toward meeting it.
Self-Hosting LLMs on Isolated GPU Infrastructure to Stay in Scope
The only durable answer for CUI-touching AI right now is self-hosting an open-weight model on hardware you control, isolated from any external network, inside your existing CMMC boundary. Waiting for a GPU cloud to announce FedRAMP authorization isn't a plan, it's a bet on a timeline nobody's committed to.
Multi-Tenant Cloud vs Dedicated, Boundary-Controlled Capacity
A shared, multi-tenant GPU instance, the kind most commercial neoclouds sell by default, puts your inference workload on hardware other tenants also touch, with a hypervisor and network path outside your CMMC boundary. That's disqualifying the moment CUI is involved, independent of the provider's other certifications. What you need instead is dedicated capacity: hardware provisioned to you alone, on a network segment you control, with no other tenant sharing the box.
This is the same distinction that governs HIPAA, PCI-DSS, and ITAR workloads, and it applies just as directly here: dedicated, single-tenant hardware is a prerequisite for CUI, not a nice-to-have. If your team is weighing whether a given workload can run on shared infrastructure at all, our HIPAA-compliant GPU cloud guide covers the same self-hosting-versus-authorization-chain tradeoff for a different regulated data type, and the logic transfers almost one to one.
Network Segmentation, Confidential Computing, and No External Connectivity
Once you have dedicated hardware, the network boundary is what actually keeps you in scope. That means the inference cluster sits on a segmented network with no unapproved outbound connectivity, no calling home to a vendor's telemetry endpoint, no default model-update check-in, nothing that lets data leave the enclave without an explicit, logged, approved path.
For teams that want hardware-level assurance on top of network segmentation, NVIDIA's Confidential Computing mode encrypts GPU VRAM during computation and adds remote attestation, letting you cryptographically verify the hardware and firmware state before trusting it with CUI, available on H100, H200, and B200. Our confidential GPU computing guide walks through the attestation flow and KMS integration in detail. It's not a CMMC requirement by name, but it's a real technical control that directly supports the "verify the environment before trusting it" posture CSA's guidance describes, and it closes the gap between "the network is isolated" and "the compute itself can't be read by a compromised hypervisor."
Fully air-gapped deployment is the extreme end of this spectrum and it's already proven at defense scale: Iternal Technologies and Intel deployed AirgapAI, a fully offline LLM system, for U.S. military use, processing an 11-million-word document set in about two hours and generating roughly 63,953 responses with zero internet connectivity throughout (Iternal Technologies, AirgapAI case study). That's the reference point for what "no external connectivity" looks like when it's done for real rather than described in a policy document.
What Open-Weight Models Fit the Hardware You Can Actually Isolate
Size the model to what you can physically isolate, not to whatever's biggest on a leaderboard. A quantized 70B-class model, Llama, Qwen, or Mistral, runs comfortably on a single 80GB H100 or H200 and covers the bulk of document summarization, technical-writing assistance, and internal Q&A use cases a defense contractor actually needs. Our best open-source LLMs to self-host guide breaks down which models fit which VRAM tier, and our AWQ quantization guide covers cutting memory footprint by roughly half with minimal quality loss, useful if your isolated hardware budget is tighter than the model you want.
Don't reach for a 400B+ frontier model unless the mission genuinely requires it. Every additional GPU in the cluster is another piece of hardware that has to sit inside your boundary, get inventoried in your SSP, and pass the same physical and logical access controls as everything else. A right-sized 70B deployment you can fully account for beats an oversized frontier model you can't.
Building a CMMC-Aligned AI Deployment Checklist
Treat the AI system as an asset like any other in your environment. It goes in the same documentation, the same access reviews, the same audit trail. Nothing about it is exempt.
SSP, POA&M, and SPRS: Documenting the AI System Like Any Other Asset
Your System Security Plan needs an entry for the AI system: what it is, which security functions it performs, how access is managed, and where it sits on the network diagram. If it doesn't yet meet every applicable control, that gap belongs in your Plan of Action and Milestones with a real remediation date, not left undocumented. And the annual affirmation you submit to SPRS under Phase 1 needs to reflect the AI system's actual state, not an assumption that it's out of scope because it's "just a chatbot."
Logging, RBAC, and MFA for the Inference Layer
Every prompt and every response that touches CUI needs to be logged, the same audit trail requirement that applies to any other system handling controlled data. Access to the inference endpoint itself should be role-based, limited to personnel who actually need it for their work, and gated behind MFA. This isn't an AI-specific requirement, it's the standard identification-and-authentication control family in NIST 800-171 applied to a new kind of endpoint.
Questions to Ask a GPU Vendor Before Anything Touches CUI
Ask these directly, in writing, before a single CUI document reaches the deployment:
- Is this dedicated, single-tenant hardware, or a shared multi-tenant instance? Multi-tenant is disqualifying for CUI regardless of other certifications.
- Does the vendor hold an active FedRAMP ATO at Moderate or above? "Working toward it" and "authorized" are not the same answer, and as of mid-2026 no GPU-specific neocloud has the former as a completed status.
- Can the deployment run with zero unapproved outbound connectivity, no vendor telemetry, no default check-in calls?
- Does the vendor support hardware-level attestation (NVIDIA CC mode or equivalent) if your risk posture calls for it?
- Who has physical and administrative access to the hardware, and is that access documented anywhere you can audit?
If a vendor can't answer all five clearly, the honest answer is that their platform isn't ready for your CUI workload yet, whatever else is on their compliance page. That said, not every workload on your team touches CUI. For unclassified R&D, model evaluation on public data, and internal prototyping, Spheron pools GPU capacity from 5+ providers through a single API, a legitimate, fast option for that category of work (Spheron overview docs). It is not FedRAMP authorized and isn't a destination for CUI, a distinction worth being explicit about rather than letting a sales conversation blur it.
If your AI workload genuinely never touches CUI, Spheron's pooled GPU capacity is a fast option for R&D and model evaluation. If it does touch CUI, this guide's whole point stands: self-host on isolated, boundary-controlled hardware, and don't take a vendor's word for a certification it doesn't hold.
Frequently Asked Questions
Not on the original schedule. DoD suspended the transition to CMMC Phase 2, which was set to require C3PAO-certified Level 2 assessments starting November 10, 2026, in a July 13, 2026 announcement. A 60-day CMMC Reform Task Force is reviewing the program, with a public RFI comment window open through August 14, 2026. Phase 1, which requires annual Level 1 self-assessments and SPRS affirmations, remains in effect and is not affected by the pause.
No. Pasting CUI into a commercial AI tool like ChatGPT or Copilot is a spillage event and a CMMC scope violation, regardless of whether the device belongs to the contractor. Commercial AI platforms don't carry the FedRAMP Moderate authorization DFARS 252.204-7012 requires for systems that touch CUI, and once the text leaves the browser, the organization has no way to pull it back.
Yes. Under 32 CFR Part 170, any cloud service that processes, stores, or transmits CUI, an AI or GPU platform included, is treated as a Cloud Service Provider and must carry FedRAMP Moderate authorization or equivalent. As of mid-2026, no neocloud or GPU-specific IaaS provider holds an active FedRAMP ATO, which is why most contractors handling CUI self-host on isolated, boundary-controlled hardware instead.
NIST SP 800-171 defines 110 security requirements organized into 14 control families, covering access control, audit and accountability, configuration management, identification and authentication, incident response, and more. Defense contractors have technically been obligated to implement all 110 since DFARS 252.204-7012 took effect in 2017; CMMC Level 2 formalizes third-party verification of that same baseline.
Anything in the Llama, Qwen, or Mistral families that fits your isolated hardware's VRAM budget once quantized. A quantized 70B-class model runs comfortably on a single 80GB H100 or H200; larger 400B+ class models need multi-GPU nodes. The constraint isn't model choice, it's that the hardware has to sit inside your boundary with no external connectivity, so size the deployment to what you can actually isolate, not to whatever fits a shared cloud instance.
