Yes, you can use Codex CLI with a local 125B model: Strata v0.1.39 added the OpenAI Responses API that Codex needs, and everything runs on your own machine.
Versions in this guide: at the time of publishing (October 5, 2026), the current version is Strata v0.1.39, released on October 4, and the pairing the project tested is with Codex CLI 0.160.0. Strata ships releases almost daily: check the latest release notes before you follow these steps.
Strata is an open source inference engine (MIT) that runs Qwen3.8-Flash-Next on a single consumer GPU, and since this version it exposes POST /v1/responses, the protocol Codex uses with custom providers. Here are the exact config and the hardware it actually takes.
What is Strata?
Strata is an inference engine built to run one model, Qwen3.8-Flash-Next (125B parameters, mixture-of-experts), on a consumer GPU. It doesn’t try to fit the whole model in VRAM; it splits the work instead:
- GPU: attention, routers, the KV cache, and an expert cache filled with the most-used experts. It adapts while you work.
- RAM: all 24,576 experts. The CPU computes whichever ones aren’t on the GPU, in parallel with it.
- SSD: a 28.8 GB n-gram table, read a few rows per token.
It serves three APIs on http://127.0.0.1:8080: OpenAI Chat Completions, Anthropic Messages and, since v0.1.39, the OpenAI Responses API. Strata itself is MIT; the model weights carry their own licenses.
You pick a variant at install:
- Q2_0 (66 GB download): fastest.
- IQ2_XS (68 GB): slightly better quality.
- IQ3_XXS (76 GB): best quality, slowest.
- Coder: an expert-pruned release tuned on code and agent data. Its authors report 91.3% of the full model’s SWE-bench Verified score.
- Swift 1.5: a fine-tune that reasons with fewer tokens.
What hardware do you need to run a 125B local AI?
For the full model you need 12 GB of VRAM and 64 GB of RAM. This is a desktop story, not a laptop one.
| Requirement (Strata docs, v0.1.39) | |
|---|---|
| GPU | NVIDIA RTX 20/30/40/50 with 12 GB VRAM or more (8 GB runs, slowly), or supported AMD Radeon cards |
| RAM | 64 GB recommended |
| CPU | x86-64 with AVX2 (AVX-512 is a bit faster) |
| Disk | 70–80 GB for the model plus ~6 GB for the draft layer; NVMe strongly recommended |
| OS | Windows 10/11 or Linux |
What about 32 GB of RAM? The full model doesn’t fit, but the Coder variant does: its experts take about 23 GB of RAM. A 32 GB PC with a 24 GB GPU can also run Q2_0 and IQ2_XS in the low-RAM mode, which copies into RAM only the experts the GPU doesn’t hold.
How fast is it? These are the project’s own figures, from a single machine: RTX 5070 12 GB, Ryzen 5 7600, 64 GB DDR5.
- Q2_0 output: about 93 tokens/s at 4K context and about 74 tokens/s at 128K.
- Prompt reading: about 1,300 tokens/s at 4K. Time to first token is roughly 4 s at 4K, 25 s at 32K, and under 2 minutes at 128K.
- Coder output: about 55 tokens/s at 4K.
- Other GPUs: the docs give figures for cards such as the RTX 3090, but label them as estimates (±20%), not measurements.
- The project’s rule of thumb: more VRAM beats a faster GPU. Each extra GB holds about 700 more experts the CPU no longer has to compute.
If the VRAM-versus-parameters math is new to you, we walked through it with another model in Kolibri de Aleph Alpha.
How do you install Strata on Windows and Linux?
On Windows, double-click START-HERE.bat. You only need an NVIDIA driver (version 580 or newer). Everything else installs itself: Python if you have none, the CUDA libraries and the engine.
The first run asks four questions:
- Which model?
- Which size?
- How much context?
- Images on or off?
It then downloads the model; an interrupted download resumes where it stopped. When it finishes, it opens the chat page at http://127.0.0.1:8080. Later starts skip setup and take 30–90 seconds to load the experts into RAM.
On Linux, run ./setup.sh. It asks the same questions and starts the same way.
Useful variants:
START-HERE.bat --setup --family coder installs the Coder variant
START-HERE.bat --model IQ2_XS --context 32768 --vision yes --yes no questions
To update, run UPDATE.bat, or ./update.sh on Linux.
How do you use Codex CLI with a local model?
Add a custom provider to ~/.codex/config.toml. On Windows the file is %USERPROFILE%\.codex\config.toml. This is the config from Strata’s documentation:
model = "strata" # any name; Strata answers with its model
model_provider = "strata"
model_context_window = 32768 # the context you chose in setup: Codex compacts before it gets there
show_raw_agent_reasoning = true # show the model's thinking (it writes no summaries)
# model_reasoning_effort = "medium" # none, low, medium or high; default: the model's (high)
[model_providers.strata]
name = "Strata (local)"
base_url = "http://127.0.0.1:8080/v1"
wire_api = "responses"
stream_idle_timeout_ms = 600000 # a first, long prompt can take minutes to read
# env_key = "STRATA_API_KEY" # only if the server has an api_key: the variable holding it
Four details matter:
- It must be your user-level config. OpenAI’s reference says Codex ignores
model_providerandmodel_providersin a project’s.codex/config.toml. model_context_windowmust match the context you picked in Strata’s setup, so Codex compacts the conversation before Strata refuses it.wire_api = "responses"is the only protocol Codex supports for custom providers. That’s why Strata needed the new endpoint.- Raise the idle timeout. Codex’s default
stream_idle_timeout_msis 300000 (5 minutes), and a long first prompt can exceed it.
Then run codex, or codex exec "...", as usual. The warning Model metadata for ... not found is expected with a local model name.
On Windows, in the project’s test, Codex’s sandbox rejected every shell command until Codex was started with -c 'windows.sandbox="unelevated"'. That’s a Codex setting, not a Strata one.
If you expose Strata beyond your own machine, set a key first: add "api_key" to strata-<model>.json and uncomment env_key in the Codex config. By default the server only listens on 127.0.0.1.
How fast is Codex CLI with a local model?
The first turn is slow and later turns are fast, because Strata caches the conversation. In the project’s test with Codex 0.160.0, using Q2_0 on the RTX 5070:
- Codex’s first prompt was 9,443 tokens of instructions and tool descriptions, read in 10 seconds.
- In a tool loop, each later turn reused about 96% of the prompt from the cache and read only the new part, in 1–2 seconds.
These numbers are self-reported by the project, from a single PC.
What doesn’t work yet?
At the time of publishing, with v0.1.39:
- No server-side state.
previous_response_idisn’t supported. Codex doesn’t need it, because it resends the conversation every turn. - No hosted tools. Tools such as
web_searchare left out, because the model can’t run them. Function tools do work. tool_choicecan’t force a tool. The model always chooses.- No reasoning summaries. The model doesn’t write them, which is why the config turns on
show_raw_agent_reasoning. - One request at a time by default.
"parallel": Ndecodes several conversations together, but on a 12 GB card it costs speed: with four requests, decoding dropped to 63.1 tokens/s against 70.7 with one. - Older and other hardware is experimental. Pascal/Volta cards, Intel Arc (Linux only, untested on Arc in this release) and CPUs without AVX2 have opt-in paths that community members wrote and measured on their own machines. The maintainer has none of that hardware.
Strata or Ollama for Codex?
If your model already runs in Ollama or LM Studio, use that: Codex has built-in providers for both (--oss, oss_provider). Strata is a different bet: it exists to run one specific 125B model on a single consumer GPU, and it plugs into Codex as a custom provider. If you’re choosing between local engines, we compared the three main ones in LM Studio, Ollama o llama.cpp: qué corre de verdad en tu máquina. And if you already route Codex to other providers through a gateway, 9router covers that path.
Free Codex? What it costs to use Codex CLI with a local model
Strata is free and MIT-licensed. Each model’s weights have their own license: the Coder is Apache-2.0 per its card, and Swift 1.5 uses the Swift Open License 1.0. The cost moves to your hardware and your power bill: Codex’s requests go to your own machine instead of a metered API.
What about you? What would you rather run Codex against: a cloud model or one on your own machine?
- En una API en la nube
- En un modelo local
- En ambos, según la tarea
- Prefiero otra opción (cuéntanos cuál)
Comment below or, if this article reached you by email, reply directly to the email: your reply gets published here.
Related: Claude Code vs Codex: cómo correr los dos con tus propias claves