The Meta Llama API shutdown became official on July 6, 2026: Meta's hosted endpoint at api.llama.com stopped answering requests that day. If you had code calling it, that code is returning a sunset response now, not a completion. Llama 4 itself, the mixture-of-experts family behind Scout (109B total, 17B active) and Maverick (400B total, 17B active), didn't disappear. What Meta shut down was the convenience layer that let you call those models by API key without running any hardware.
The interesting part is what Meta pointed developers toward instead, and how much of it has already come apart. Meta's own migration guidance named third-party hosts, Groq chief among them, as the fallback for hosted Llama 4 access. Groq deprecated Llama 4 Maverick on March 9, 2026 and Llama 4 Scout on July 17, 2026, eleven days after the Llama API itself went dark. The official off-ramp is shrinking under the people who took it. This post covers exactly what shut down, what didn't, what Meta charged versus what the remaining resale APIs charge for Llama 4 right now, and the GPU math for the path that doesn't depend on anyone else's roadmap: renting your own hardware and self-hosting it.
TL;DR: Is the Meta Llama API Shutdown Permanent, and What Should You Do?
- Shutdown date: Meta's Llama API at
api.llama.comretired on 6 July 2026; requests now return a sunset response. - What survives: Llama 4 Scout and Maverick weights stay open-weight and downloadable under Meta's license.
- Migration path thinning: Groq deprecated Llama 4 Maverick on 9 March 2026 and Llama 4 Scout on 17 July 2026.
- Cheapest resale: DeepInfra prices Llama 4 Scout at $0.10/M input and $0.30/M output tokens.
- Self-hosting floor: Scout runs on one H100 80GB; Spheron prices H100 on-demand at $2.64/hr, as of 28 Sep 2026.
- Verdict: With Meta's hosted option and its named Groq fallback both gone, self-hosting is the only path with a stable floor. Compare current GPU rates.
What Actually Shut Down on July 6, 2026 (and What Didn't)
Two different things happened to Llama 4 this year, and conflating them is what leads to the "wait, is Llama 4 dead?" confusion showing up in developer forums. One is permanent. The other is a service Meta chose to stop running.
The Exact Cutoff: api.llama.com Retired, Requests Return a Sunset Response
Meta's Llama API, its own first-party hosted endpoint for calling Llama 4 by API key, retired on July 6, 2026. Requests to api.llama.com no longer complete; they return a sunset response instead. Promptfoo's provider documentation confirms this and notes that the llamaapi: provider identifier is still registered in third-party SDKs for backward compatibility, but that registration doesn't make the endpoint work. A library that still lists the provider isn't evidence the service is still live; it's just a library that hasn't been pruned yet.
Promptfoo's own migration guidance splits into three paths once the Llama API is gone: Meta's separate Model API for Meta-hosted inference outside the retired endpoint, third-party hosts including AWS Bedrock, Together AI, or Groq for hosted Llama access, and local deployment via Ollama, llama.cpp, or vLLM. That third path, running the model yourself, is the one this post spends the most time on, because it's the only one that doesn't depend on a vendor's continued willingness to serve a specific model.
If you're running llama.cpp already and just need the multi-GPU and GGUF specifics, our llama.cpp server deployment guide covers that setup directly.
What Survives: Llama 4 Weights Stay Open and Downloadable
Nothing about the shutdown touches the weights. Llama 4 Scout and Maverick remain open-weight, downloadable from Hugging Face under Meta's community license, exactly as they were before July 6. This answers the "is Llama 4 free" question directly: the model itself is free to download and run on your own hardware. What disappeared was Meta's own hosted convenience layer, the option to skip owning any infrastructure and just pay Meta per API call. Free-to-download and free-to-call were never the same guarantee, and only the second one went away.
That distinction matters more this year than it would have a year ago. Meta shipped no new open-weight Llama release through the first half of 2026, and its frontier attention moved instead to the closed, paid Muse Spark API. Llama 4 Scout and Maverick, both from April 2025, remain the newest generation of open Llama weights Meta has shipped. Our Muse Spark pricing versus self-hosted breakdown covers that pivot in detail; the Llama API shutdown reads less like an isolated event and more like the second half of one strategic move away from giving away free, hosted access to frontier-adjacent models.
Why Meta's Own Migration Advice Is Already Half-Broken
Here's the part every other post on this shutdown misses. Meta's migration guidance treats Groq, AWS Bedrock, Together AI, and DeepInfra as a stable, interchangeable set of alternatives. They aren't stable, and the gap is showing up in the exact provider Meta's own guidance leads with.
Groq Deprecated Llama 4 Maverick and Then Scout, Eleven Days After the API Itself Died
Groq's model deprecation log shows Llama 4 Maverick (meta-llama/llama-4-maverick-17b-128e-instruct) deprecated on March 9, 2026, with Groq recommending openai/gpt-oss-120b as the replacement. Then Llama 4 Scout (meta-llama/llama-4-scout-17b-16e-instruct) followed on July 17, 2026, eleven days after api.llama.com itself went dark, with Groq recommending openai/gpt-oss-120b or qwen/qwen3.6-27b in its place.
Read that sequence in order. Meta's own migration documentation names Groq as a hosted fallback for Llama access. By the time the Llama API had been dead for less than two weeks, Groq had already deprecated both Llama 4 variants and was steering its own customers toward OpenAI's and Alibaba's open models instead of Meta's. A reader following Meta's official advice today would land on Groq, try to call Llama 4, and get pointed at a completely different model family. If you want Groq's current price ladder for what it does still serve (and the same deprecation reality reflected in real numbers), our Groq pricing breakdown covers its live catalog, which by mid-2026 carries no Llama 4 model at all.
What's Actually Left: AWS Bedrock, Together AI, DeepInfra, Novita
Strip Groq out and the field of active Llama 4 hosts is smaller than the "four major alternatives" framing suggests. Artificial Analysis's own Llama 4 Maverick provider benchmark lists six active hosts as of its latest measurement: DeepInfra, Amazon Bedrock, Azure, Novita, Parasail, and Together AI. Groq does not appear on that list. That's independent, third-party confirmation of the same gap Groq's own deprecation log shows.
This isn't a reason to panic about Llama 4 disappearing entirely, since real capacity does remain across DeepInfra, Bedrock, Together AI, and Novita. It's a reason not to build a migration plan around a provider list that was already stale by the time the shutdown notice went out. The next section prices out what those remaining hosts actually charge.
Llama 4 API Pricing Now: What Novita, DeepInfra, and the Rest Charge
What Meta Charged Before Retiring Its Own API
There's no official Meta price list to hold up against the third-party rates below, and that absence is itself the finding. Every URL that used to point at Meta's own Llama API now redirects somewhere else: llama.developer.meta.com and developer.meta.com/ai/products/llama-api/ both forward to dev.meta.ai, a page that promotes Meta's separate Model API and Muse products but carries no Llama API pricing information at all. When it retired the service on July 6, 2026, it didn't leave a stable, citable per-token price list behind for anyone comparing "official vs. third-party" the way you'd compare it for a still-live API. Every dollar figure available for Llama 4 API access today, the table below included, comes from a third party. That's not a gap in this post's research; it's the actual state of the market.
Llama 4 Scout and Maverick Price Table (Per Million Tokens)
Here's the current per-million-token pricing across the providers still serving Llama 4 Scout and Maverick, gathered directly from each provider's own pricing page.
| Provider | Model | Input ($/M tokens) | Output ($/M tokens) |
|---|---|---|---|
| DeepInfra | Llama 4 Scout | $0.10 | $0.30 |
| Novita | Llama 4 Scout Instruct | $0.18 | $0.59 |
| Novita | Llama 4 Maverick Instruct | $0.27 | $0.85 |
| DeepInfra (FP8, per Artificial Analysis) | Llama 4 Maverick | $0.20 | $0.80 |
DeepInfra's own pricing page lists Llama 4 Scout at $0.10 per million input tokens and $0.30 per million output tokens, the cheapest headline Scout rate in this table. Novita's pricing page lists Llama 4 Scout Instruct at $0.18/$0.59 and Llama 4 Maverick Instruct at $0.27/$0.85 per million input/output tokens. For Maverick specifically, Artificial Analysis's provider comparison shows DeepInfra's FP8 listing at $0.20 input / $0.80 output per million tokens, the cheapest Maverick input rate among the six providers it benchmarks.
Converted down: at DeepInfra's Scout rate, 1,000 input tokens cost $0.0001 and 1,000 output tokens cost $0.0003. At Novita's Maverick rate, 1,000 input tokens cost $0.00027 and 1,000 output tokens cost $0.00085. Those per-thousand numbers are why per-token pricing feels free right up until a real production workload multiplies them by millions of daily requests, at which point the comparison to a flat hourly GPU rate stops being abstract.
Pricing fluctuates based on provider capacity and can change without notice. DeepInfra and Novita rates above reflect their published pricing pages as of this post's writing; check DeepInfra's pricing page and Novita's pricing page directly for current rates before budgeting a production workload. For a broader look at DeepInfra's dedicated-GPU tier alongside its per-token rates, see our DeepInfra pricing breakdown.
The Self-Hosting Path: GPU Sizing for Llama 4 With No Meta Fallback
With the official hosted option gone and its named fallback provider already deprecating both Llama 4 variants, self-hosting stops being one option among several and becomes the option with a stable floor. Nobody can deprecate a model you're running on your own rented GPU.
Scout vs Maverick, Minimum Viable GPU Tier
Scout and Maverick have very different footprints, both mixture-of-experts models where every expert's weights sit resident in memory whether or not a given token routes through them, so parameter count alone understates the sizing problem.
| Model | Total / active params | INT4 weight footprint | Minimum viable GPU tier |
|---|---|---|---|
| Llama 4 Scout | 109B / 17B | ~55GB | 1x H100 80GB |
| Llama 4 Maverick | 400B / 17B | ~200GB | 4x H100 80GB (INT4); 8x H100 DGX host at FP16 per Meta's own guidance |
Scout's roughly 55GB INT4 footprint leaves about 25GB of headroom on a single 80GB H100 for KV cache, activations, and framework overhead, enough for moderate context lengths in the tens of thousands of tokens. Maverick's 200GB INT4 footprint exceeds any single card on the market, so a 4x H100 80GB node (320GB) is the realistic floor for a QLoRA-friendly deployment, with Meta's own guidance recommending a full 8x H100 DGX host at FP16 for production. Our LLM VRAM calculator for Llama 4 walks through the exact FP16/INT8/INT4 math and where KV cache growth eats into that headroom at longer context windows, and deploying Llama 4 on GPU cloud with vLLM covers the actual serving setup once you've picked a tier.
If you also need to customize the model rather than just serve it, that path used to run through Meta's hosted fine-tuning options for some frontier models; with no Llama API left to lean on, fine-tuning Llama 4 on rented GPUs is now the only route, and it's a viable one: Scout fits Unsloth's 4-bit QLoRA path on a single 80GB card.
Cost Comparison: Resale API Tokens vs Renting Your Own GPU
The honest version of this comparison depends entirely on volume, not on which option is cheaper in the abstract. At low, bursty token counts, a per-token resale API from DeepInfra or Novita can beat renting a GPU that bills whether or not it's generating tokens. At sustained volume, the math flips, because a rented GPU's cost is fixed per hour regardless of how many tokens it produces, while a resale API's cost scales linearly with every token you send.
| Approach | Model | Cost structure | Rate | Scales with token volume? |
|---|---|---|---|---|
| Resale API (DeepInfra) | Llama 4 Scout | Per-token | $0.10/M input, $0.30/M output | Yes, linearly |
| Resale API (Novita) | Llama 4 Maverick | Per-token | $0.27/M input, $0.85/M output | Yes, linearly |
| Self-hosted GPU (Spheron H100, on-demand) | Llama 4 Scout, 1x H100 80GB | Flat hourly | $2.64/hr | No, fixed regardless of tokens sent |
| Self-hosted GPU (Spheron H100, on-demand) | Llama 4 Maverick, 4x H100 80GB node | Flat hourly, node total | $10.56/hr ($2.64/hr each) | No, fixed regardless of tokens sent |
Spheron is a GPU marketplace that aggregates availability across 5+ providers, offering bare-metal H100, H200, B200, and B300 instances with per-minute billing after a 20-minute minimum, so a Scout deployment costs only for the minutes it actually runs rather than a fixed monthly commitment. As of 28 Sep 2026, Spheron prices H100 on-demand at $2.64/hr and spot at $2.21/hr, with H200 on-demand at $4.80/hr and spot at $2.93/hr. Bare-metal access with full root is what a self-hosted vLLM deployment for Llama 4 actually needs, and Spheron's lack of lock-in contracts means a team can run a self-hosted Scout deployment against a resale API's real invoice for a week before committing either way.
Pricing fluctuates based on GPU availability. Spheron rates above are live as of 28 Sep 2026; other providers reflect their most recent published rates and may have changed. Check current GPU pricing → for live rates.
That said, self-hosting isn't free of tradeoffs even once the API disappears. Renting a GPU takes on model ops, the serving stack, scaling, and monitoring, that a resale API abstracts away entirely, and spot instances can be reclaimed without notice, which matters for anything production-facing. A workload sending a few thousand requests a day is often genuinely cheaper on DeepInfra's $0.10/$0.30 Scout rate than on a dedicated GPU sitting mostly idle. The crossover point is a function of daily token volume, not a fixed rule, so run both numbers against your own traffic before deciding. A bare-metal H100 GPU rental is the exact instance type both the Scout and Maverick sizing tables above assume, and Spheron's own docs cover the deployment steps if you're provisioning for the first time.
Migrating Off the Llama API: A Practical Checklist
If you had production code pointed at api.llama.com, here's the order to work through the migration:
- Confirm the outage isn't transient. Check for a sunset response rather than a timeout or rate-limit error; the retirement is permanent, not a temporary capacity issue.
- Estimate your real daily token volume, split by input and output, since resale API pricing and self-hosted GPU costs scale differently against each.
- Price the resale APIs still standing against your volume: DeepInfra's $0.10/$0.30 Scout rate or Novita's $0.27/$0.85 Maverick rate from the table above, using each provider's own current pricing page rather than a cached number.
- Don't default to Groq without checking its current model list first. It deprecated both Llama 4 variants within months of the API shutdown; verify what it actually serves today before building a migration plan around it.
- Size the self-hosting option in parallel, even if you land on a resale API for now. Scout's single-H100 floor and Maverick's 4x H100 floor (see the table above) are cheap enough to test against real traffic before committing to either path long-term.
- If self-hosting, pick a serving stack. A direct vLLM deployment for the fastest path to a working endpoint, or Llama Stack if you want Meta's own production framework for inference, agents, and RAG bundled together rather than assembling it yourself.
- Re-check pricing before committing to volume. Every rate in this post, ours and every third party's, can move; confirm current numbers on the provider's own page before signing anything that assumes today's price holds.
If Llama 4 turns out to be more model than your workload needs, our GPU requirements cheat sheet tables out VRAM and live cost for 18 other models, in case a smaller checkpoint clears the bar just as well for less hardware.
The Llama API's shutdown didn't touch the weights, only the convenience of not running your own hardware. Renting a bare-metal H100 to self-host Llama 4 Scout costs $2.64/hr on-demand as of 28 Sep 2026, with no dependency on a third party's deprecation schedule.
Frequently Asked Questions
The weights are. Meta still distributes Llama 4 Scout and Maverick as open-weight downloads under its community license, and that didn't change on July 6, 2026. What shut down was Meta's own hosted API at api.llama.com, the service that let you call Llama 4 by API key without running any hardware yourself. Free-to-download and free-to-call are different things, and only the second one went away.
Among the providers still serving Llama 4 as of this writing, DeepInfra prices Scout at $0.10 per million input tokens and $0.30 per million output tokens, the lowest headline rate in this post's table. Novita prices Scout at $0.18/$0.59 and Maverick at $0.27/$0.85 per million tokens. For Maverick specifically, DeepInfra's FP8 listing at $0.20/$0.80 input/output is the cheapest input rate among the six providers Artificial Analysis benchmarks. Rates move with provider capacity, so check each provider's own pricing page before committing to a volume estimate.
At DeepInfra's Scout rate of $0.10/M input and $0.30/M output, 1,000 input tokens cost $0.0001 and 1,000 output tokens cost $0.0003. At Novita's Maverick rate of $0.27/M input and $0.85/M output, 1,000 input tokens cost $0.00027 and 1,000 output tokens cost $0.00085. These per-thousand figures look trivial in isolation; the ones that add up are the daily and monthly totals once you multiply by real request volume, which is exactly where a rented GPU's flat hourly rate starts to compete.
Llama 4 Scout (109B total, 17B active) fits on a single H100 80GB at INT4 quantization, with roughly 55GB of weights leaving headroom for KV cache at moderate context lengths. Llama 4 Maverick (400B total, 17B active) needs at least 4x H100 80GB at INT4, or Meta's own recommended 8x H100 DGX host at FP16. Neither model requires exotic hardware; both run on GPU tiers a marketplace like Spheron rents by the hour.






