Groq shut down llama-3.1-8b-instant and llama-3.3-70b-versatile for free and developer-tier usage on August 16, 2026. If an app, bot, workflow, or test suite started failing on Groq today, check whether it still sends one of those model IDs.
Direct answer
The fix is to replace the retired Groq model ID, then retest your prompts before pushing the change to production. Groq recommends openai/gpt-oss-20b for llama-3.1-8b-instant, and openai/gpt-oss-120b or qwen/qwen3.6-27b for llama-3.3-70b-versatile.
| Old Groq model ID | Shutdown date | Groq’s recommended replacement |
|---|---|---|
llama-3.1-8b-instant |
August 16, 2026 | openai/gpt-oss-20b |
llama-3.3-70b-versatile |
August 16, 2026 | openai/gpt-oss-120b or qwen/qwen3.6-27b |
Do not assume the replacement will produce identical answers. The fastest safe path is to change the ID in a staging branch, run the prompts or evaluations that matter, compare cost and output length, then ship the small config update.
What Groq has confirmed
Groq’s deprecation page confirms these points:
- The affected model IDs are
llama-3.1-8b-instantandllama-3.3-70b-versatile. - The shutdown date is August 16, 2026.
- Groq says it emailed affected users on June 17, 2026.
- The change applies to free and developer-tier usage.
- Enterprise customers with a committed-spend contract are not affected by this specific deprecation.
- After end of life, deprecated model IDs are no longer accessible and requests to them return errors.
That means this is not a general Groq outage. If only one workflow fails, start by searching your code and environment variables for the old model ID.
What remains unknown
Groq does not know how your app uses the model. These points still depend on your own setup:
- Whether the model ID is hard-coded in app code, a dashboard setting, a
.envfile, or a third-party tool. - Whether your prompt needs changes because the replacement model behaves differently.
- Whether your provider wrapper still exposes the old model as a selectable option.
- Whether your account tier or enterprise contract changes the deadline.
- Whether your tests cover the failed path.
If you use Groq through another product, check that product’s provider catalog too. Some tools keep compatibility rows after an upstream provider has announced a shutdown.
Quick migration checklist
Use this order before changing more code than needed:
- Search for
llama-3.1-8b-instantandllama-3.3-70b-versatilein code, config, secrets templates, examples, docs, and saved dashboard settings. - Replace only the model ID first.
- Run your most common prompts against the replacement model.
- Compare response shape if your app expects JSON, tool calls, fixed labels, or short answers.
- Check max output tokens and context window assumptions.
- Compare Groq’s published pricing for your old and new model choices.
- Ship the config change only after the important paths pass.
For broader tool selection, see how to compare AI tools before you try them. If your team is also checking plan limits, AI tool pricing explains the common credit and subscription traps.
What to change in code
If your app uses a single model constant, the change may be only one line:
const GROQ_MODEL = "openai/gpt-oss-20b";
For a larger app, put the model ID in one setting and avoid scattering it through handlers, tests, and scripts. That makes the next provider change a config update instead of a repo-wide edit.
If you call Groq through OpenAI-compatible clients, the base URL does not need to change just because the model ID changed. Groq’s OpenAI compatibility docs still show the https://api.groq.com/openai/v1 base URL. Update the model value and retest the same request path.
Which replacement should you try first?
The closest first test depends on what you used:
| If you used | Try first | Why |
|---|---|---|
llama-3.1-8b-instant for cheap, fast chat |
openai/gpt-oss-20b |
Groq names it as the replacement and lists higher token speed than the old 8B model. |
llama-3.3-70b-versatile for larger text tasks |
openai/gpt-oss-120b |
Groq names it first among replacements and lists a lower input price than the old 70B model. |
llama-3.3-70b-versatile with multimodal needs |
qwen/qwen3.6-27b |
Groq lists Qwen 3.6 27B as another replacement and marks a max file size in the models table. |
Treat this as a test order, not a universal ranking. A cheaper or faster model can still fail your exact prompt format.
Troubleshooting failed Groq calls
If the app is still broken after changing the model ID, check these common causes:
- The old model ID also exists in tests, worker jobs, examples, or fallback code.
- A third-party integration has its own provider model setting.
- Cached deployment variables still point to the retired ID.
- The replacement model returns longer output than your parser expects.
- Your code treats every 400, 403, or 404 response as an auth problem.
- Your account does not have access to the replacement model you selected.
Groq’s error-code docs separate client errors such as bad request, unauthorized, forbidden, and not found. Read the status and message before rotating keys or changing unrelated auth code.
Official sources checked
Groq’s model deprecation page is the primary source for the August 16, 2026 shutdown date, affected model IDs, tier note, and recommended replacements.
Groq’s supported models page lists current model IDs, token speed, developer-plan rate limits, context windows, max completion tokens, prices, and the qwen/qwen3.6-27b preview row.
Groq’s OpenAI compatibility page confirms the OpenAI-compatible base URL pattern. Groq’s API error codes page explains the status codes to check when a request fails.
FAQ
Why did my Groq Llama API call stop working?
The most likely same-day cause is that your request still uses llama-3.1-8b-instant or llama-3.3-70b-versatile. Groq lists August 16, 2026 as the shutdown date for those IDs on free and developer-tier usage.
What replaced llama-3.1-8b-instant on Groq?
Groq recommends openai/gpt-oss-20b as the replacement for llama-3.1-8b-instant.
What replaced llama-3.3-70b-versatile on Groq?
Groq recommends openai/gpt-oss-120b or qwen/qwen3.6-27b as replacements for llama-3.3-70b-versatile.
Does this affect Groq enterprise customers?
Groq says this specific deprecation applies to free and developer-tier usage, and that enterprise customers with a committed-spend contract are not affected.
Do I need to change my Groq API base URL?
Usually no. If you already use Groq’s OpenAI-compatible endpoint, update the model value first and retest before changing client setup.
Should I switch every prompt to the new model without testing?
No. Change the model ID in staging, run important prompts or evaluations, check structured output and costs, then deploy the config update.