Choosing the Right Codex Model in VS Code

A practical guide to GPT-5.6 Luna, Terra, Sol, GPT-5.5, and GPT-6 Astra in the Codex VS Code plugin, including when to use each model, how the reasoning bar affects behavior, and how model choice impacts token usage and cost.

In the Codex extension for Visual Studio Code, you are choosing two things:

  1. The model — its overall capability, speed, and usage rate.
  2. The reasoning-effort bar — how much computation that model should spend on the current task.

The reasoning bar is not model fine-tuning in the machine-learning sense. It temporarily changes how deeply Codex reasons. It does not retrain or permanently modify the model.

Which model should you use?

ModelBest use in VS CodeRelative usage
GPT-5.6 LunaExplanations, small edits, boilerplate, simple tests, repetitive tasksLowest
GPT-5.6 TerraNormal daily coding, debugging, tests, small-to-medium featuresLow
GPT-5.6 SolComplex features, architectural changes, difficult debugging, coordinated multi-file workMedium-high
GPT-5.5Older strong general model; useful for compatibility or comparing behaviorHigh
GPT-6 AstraThe hardest repository-wide work, ambiguous bugs, migrations, security/architecture reviews, long autonomous tasksHighest per token

OpenAI’s current guidance recommends Astra for maximum capability, Terra for the intelligence/cost balance, and Luna for cost-sensitive high-volume work. Sol is the flagship GPT-5.6 professional model. Official model catalog

My practical VS Code recommendation:

  • Set GPT-5.6 Terra + Medium as your normal default.
  • Use Luna + Low/Medium for easy mechanical work.
  • Switch to Sol + High when Terra struggles.
  • Use Astra + High only for genuinely difficult or expensive-to-get-wrong tasks.
  • Keep GPT-5.5 mainly when you prefer its behavior or need to compare an existing workflow against the newer family.

GPT-5.6 Sol is generally a better current choice than GPT-5.5: it is newer, more capable, and its published token/credit rate is lower.

What does the reasoning-effort bar do?

Depending on the selected model and extension version, you may see levels such as:

  • None/Low: Minimal planning. Fast and economical.
  • Medium: Balanced reasoning. Best default for most programming.
  • High: More analysis, repository inspection, and verification.
  • xHigh: Difficult debugging, architecture, migrations, or subtle correctness problems.
  • Max/Ultra: Maximum available effort for exceptionally difficult work.

Not every model supports every level. Astra, for example, does not support none; its documented levels start at low. Astra model guidance

Higher reasoning is not automatically better. On a simple rename or CSS adjustment, xHigh may spend time reconsidering obvious decisions, searching more files, and running additional checks without improving the result.

Does selecting a stronger model consume more tokens or usage?

Yes, generally—but there are two separate effects.

1. Different models consume your allowance at different rates

For credit-based Codex usage, the current published rates per one million input/output tokens are:

ModelInput creditsOutput credits
GPT-5.6 Luna530
GPT-5.6 Terra50300
GPT-5.6 Sol100500
GPT-5.5125750
GPT-6 Astra2501,250

Thus, for the same number of tokens, Astra consumes roughly:

  • 2.5× the credits of Sol
  • 5× Terra
  • 50× Luna

However, the actual difference per task can be smaller because a stronger model might solve the problem in fewer attempts or produce fewer output tokens. OpenAI specifically notes that Astra can sometimes complete difficult tasks with fewer output tokens than earlier models. Codex pricing and usage

2. Higher reasoning effort usually uses more tokens

Moving from medium to high, xHigh, or max can increase:

  • Internal reasoning tokens
  • Tool calls
  • Files and command output read into context
  • Time taken
  • Total usage or credits consumed

There is no fixed multiplier such as “High always costs twice Medium.” It depends on the task. A simple task might show little difference; a complex investigation might show a large difference.

Everything Codex reads also matters: your prompt, conversation history, files, AGENTS.md, terminal output, tool responses, and generated answer all contribute to usage. Large chats and unrestricted repository searches can sometimes matter more than the few words in your latest prompt.

Recommended operating pattern

Use this escalation ladder:

  1. Start with Terra + Medium.
  2. For trivial changes, drop to Luna + Low/Medium.
  3. If the task involves multiple interacting systems, use Sol + Medium/High.
  4. If Sol cannot resolve it, or the consequences of a mistake are serious, use Astra + High.
  5. Reserve xHigh, max, or ultra for problems where additional reasoning has a clear value.

Also start a fresh Codex conversation when changing to an unrelated task. Carrying a very long chat history into every request increases context usage regardless of which model you select. You can view current usage in the Codex usage dashboard; in the CLI, /status reports remaining limits.

Posted in AI

Leave a Reply

Your email address will not be published. Required fields are marked *