Library

I Tried to Fix a Small Codex Auth Problem in Caveman

cavemancodexopen-sourceauthenticationtestingcode-review

I Tried to Fix a Small Codex Auth Problem in Caveman

I wasn't trying to redesign Codex authentication. I was trying to understand why something that looked like a successful codex login could still leave the native agent using the wrong authentication path.

At first, that sounds like a small configuration problem. Login succeeds, the configuration exists, and the proxy is running, so where could the mismatch actually be?

The problem was that there wasn't just one piece of state. Codex had its own authentication state, Caveman had routing information for that state, and the native-agent startup flow had to decide which one to use. If those pieces stopped agreeing with each other, the system could look healthy while sending the request through the wrong lane.

So I stopped looking at the individual error and followed the request from the beginning.

codex login
     │
     ▼
Codex auth state changes
     │
     │
     ├── Caveman routing/config
     │
     └── Native agent startup
              │
              ▼
         SessionStart
              │
              ▼
       Which auth lane?
              │
              ▼
         Proxy routing

That was the first useful clue. The question wasn't simply "is the user logged in?" It was "which authentication state does the native agent believe it should be using when it starts?"

The state could change underneath the configuration

The important part of the investigation was codex login.

A user could change authentication mode after Caveman had already written its Codex configuration. The login command changes Codex's current authentication state, but that doesn't automatically mean every Caveman-generated routing decision changes at the same moment.

That creates a stale state problem.

Before login change

Codex
 └── auth = API key

Caveman config
 └── /w/codex/v1

       ↓

Native agent starts
       ↓
Uses API-key route

User runs:

codex login

After login change

Codex
 └── auth = ChatGPT subscription

Caveman config
 └── still contains previous auth lane

       ↓

Native agent starts
       ↓
Potential mismatch

This was the part that made the issue more interesting than a normal authentication failure. Nothing necessarily had to crash during codex login. The bad state could only become visible later, when the native agent started making requests.

The proposed fix in PR #1047 was to check the current authentication state during SessionStart and repair stale routing before continuing.

There is an important distinction here, though. That self-healing behavior is proposed work in the open PR. It was not silently folded into the merged routing fix in PR #1015. Caveman already has an explicit caveman doctor codex --fix path for detecting and repairing this kind of stale configuration, while automatically rewriting ~/.codex/config.toml from a host startup hook is a separate policy decision.

That distinction matters because "the system can repair this" and "the system should automatically rewrite the user's configuration during startup" are not quite the same thing.

Then I found another Codex problem

While following the request path, I ran into something that looked unrelated at first.

The API-key route itself was wrong.

Caveman was writing a Codex route without the /v1 segment:

/w/codex

But Codex's OpenAI Responses client constructs the final endpoint by appending /responses to its configured base URL.

So the request became:

/w/codex
    +
/responses
    ↓
/w/codex/responses

The proxy did not consider that the normal OpenAI Responses endpoint.

The route it actually needed was:

/w/codex/v1

which produces:

/w/codex/v1
    +
/responses
    ↓
/w/codex/v1/responses

That small /v1 was the difference between the request matching the OpenAI adapter and getting a 404 cave_route_not_found before the request ever reached OpenAI.

The test was making the bug harder to see

This was probably my favorite part of the investigation.

The existing conformance test was not reproducing what the real Codex client was doing. The test stub was effectively adding:

/v1/responses

while the real Codex client was adding:

/responses

That meant the test could produce a valid-looking route even though the actual configuration was missing /v1.

So the test was accidentally fixing the problem before the request reached the router.

I changed the stub to behave like the real client and then ran the test again. The broken route immediately became visible.

The route matching was then made explicit with a regression test:

/w/codex/responses
       │
       └── should NOT match OpenAI adapter

/w/codex/v1/responses
       │
       └── should match OpenAI adapter

The actual route was moved into a shared constant:

CODEX_API_KEY_ROUTE = "/w/codex/v1"

That same value is used by the provider configuration, installation journal, and the expected route reported by nativeIntegrationStatus.

This also made existing installations easier to reason about. An installation containing the old route can be reported as degraded, with the repair path pointing to:

caveman doctor codex --fix

The subscription route stays different:

/chatgpt

It does not need the /v1 suffix because it goes through a different proxy handler.

What happened to those two threads

A few weeks later, the maintainer's daily triage caught up with both threads, and the verdict was split.

The routing half was real, and it landed. PR #1015 carried the fix — commit 8bc01d6 wrote /w/codex/v1 into the shared CODEX_API_KEY_ROUTE constant — and the conformance harness was repaired so it reproduces the original 404 instead of masking it the way the old test stub did. My name went on that work as a co-author credit.

The self-healing half got a different answer. The triage declined it as written, and explicitly called it a maintainer decision, not a request for changes. A SessionStart hook restoring and reapplying ~/.codex/config.toml unasked buys convenience, not capability: caveman doctor codex --fix already repairs the drift, and doctor codex already reports it. The narrower shape suggested for the future, if it ever lands, is a hook that warns rather than writes.

I could have argued with that. I didn't, because re-reading the reasoning, I agreed with it. "The system can repair this" was proven. "The system should rewrite your configuration during startup" was never claimed, and the burden for that claim is higher.

Reading the two outcomes side by side, the difference fits in one line: a fix that rewrites user state during startup is a policy decision, while a fix that only reads configuration and prints to stdout is a patch.

I filed that lesson away. I didn't expect to test it so soon.

The second PR: a hook that was really a print statement

I wasn't planning a second investigation. I wanted to give that scope lesson one honest test. When I filed it away, it was still a theory. My next caveman PR was the experiment.

The experiment came from issue #185. Caveman's Codex integration declares a SessionStart hook in .codex/hooks.json, which sounds like the same mechanism the Claude side has. In practice, the hook was a static echo of the full-level rules. It did not resolve anything:

Codex session starts
     │
     ▼
SessionStart hook runs
     │
     ▼
static echo of the full rules   ← the entire "resolution"

Which meant, on every Codex session, CAVEMAN_DEFAULT_MODE had no effect, a repo-local .caveman/config.json was ignored, a user config defaultMode was ignored, and a configured default of off still force-injected the rules.

That last one is what made this a bug and not a missing feature. The user had asked caveman to be off, and the hook injected the rules anyway. A missing feature is a gap. A gap that inverts the user's own opt-out is a defect.

The Claude side already had a real resolver — environment variable, then repo config, then user config, then default — shared through src/hooks/caveman-config.js. Codex had a print statement. And because the rules were maintained in two places, they had drifted, exactly as cowwoc noted in the issue.

The rule I set: no second copy of the truth

PR #1139 added a small codex-sessionstart.js that requires the same shared resolver the Claude hook uses, then emits the same SKILL.md-filtered ruleset through the same loader. One source of truth, two hosts.

Everything else followed from the scope lesson:

  • the hook reads configuration and prints; it never writes
  • it fails open: any error prints nothing and exits 0, so it can never block a session
  • it never calls process.exit() after writing stdout, so the output flushes naturally
  • the command resolves from the git root — node "$(git rev-parse --show-toplevel)/.codex/codex-sessionstart.js" — so activation works from any launch directory
  • the 15 tests are hermetic, running against fixture copies, so my developer environment could not leak into the results
  • commits are DCO-signed

Two of the tests are worth naming, because of what they assert. One pipes the hook's output through another process to prove it survives intact — that is the test that makes the missing process.exit() meaningful rather than superstitious. Another pins the hand-copied fallback whitelist equal to VALID_MODES, so the degraded path and the shared resolver cannot drift apart the way the two copies of the rules did before. Each one guards a failure this story has already met: two copies drifting silently, and output that never reaches the reader.

What the review said

All 15 tests were confirmed red on the unfixed tree first. The manifest case, the maintainer noted, "fails printing the literal old echo string, which is the right reason."

Then came the list of what made the PR adoptable rather than merely plausible: resolving through the shared resolver instead of adding a second copy of the resolution order; never calling process.exit() after stdout, proven by piping the output through another process; the parity test pinning the fallback whitelist; the degraded branch resolving the full order, because ignoring a user config that says off would invert the user's own opt-out. They added exactly one thing themselves: a row in CLAUDE.md's single-source table, because the file is behavior-bearing and the table did not name it.

And then the sentence this post has been building toward:

Adopted into #1103, cherry-picked so your authorship is preserved. Thank you — this is careful work.

The commit that will land on main:

635253b feat(codex): SessionStart hook honors defaultMode
Author:    g0GobliN <[email protected]>
Committer: Claude <[email protected]>   ← the maintainer's bot

Committer and author are different on purpose, and that difference is the cherry-pick. The maintainer's tooling moved my commit into the triage branch — the move is the committer. The work, and the name on it, stayed mine.

That is also the quiet difference from the first PR. #1015 merged my diagnosis and carried my work with a co-author credit. #1139 was adopted whole, commit untouched, authorship preserved. Same repo, same reviewer — the variable that changed was scope.

Two PRs, one lesson

                     Two caveman PRs
                              │
                 ┌────────────┴────────────┐
                 │                         │
                 ▼                         ▼
        PR #1047 self-heal        PR #1139 defaultMode
                 │                         │
                 ▼                         ▼
        proposes rewriting        reads config, prints rules;
        user config at startup    never writes
                 │                         │
                 ▼                         ▼
        split: route half         adopted whole, authorship
        merged (#1015), half      preserved
        declined

The difference was not effort or care — #1047 got plenty of both. The difference was where each fix stopped. One proposed rewriting user state on startup; the other only read configuration and printed. This post called that a scope lesson once. It turned out to be the actual review line.

What it actually took

Nothing above reads like a month of work, so here is the honest inventory:

  • traced the Codex authentication state through the native-agent SessionStart flow and found the stale-lane window after codex login
  • proposed self-healing for it, and learned why that fix was a policy decision, not a patch
  • found the missing /v1 in the API-key route, fixed it, and got it merged through #1015
  • caught a conformance test stub masking the exact bug it was supposed to catch, and made it behave like the real client
  • isolated authentication tests from the local OPENAI_API_KEY so results stopped depending on the machine
  • replaced a print statement pretending to be a hook with a real resolver sharing one source of truth
  • wrote 15 hermetic tests and verified all of them red-first on the unfixed tree
  • ran the native-enable runtime suite (44/44), the TypeScript builds, and the root-level checks, twice over

Each bullet has a section above it. The straight line in the post was a lot of following the request past the point where it stopped looking like one problem.

Where it stands

As of October 1, the second fix sits as a credited commit inside the daily triage PR #1103 — adopted, CI green, not yet merged. #1139 and #185 close when that lands. The contribution graph is waiting on a merge click, not on more work.

When I started the auth investigation, I thought the finding was a routing bug. It was. But the durable finding turned out to be smaller and more useful: reviewers will argue with your fix, but they have a hard time arguing with your scope.

This is my first post about contributing to open source. If you're picking your first issue somewhere: small scope, one source of truth, no writes. It turns out that is also a pretty good recipe for getting your name kept on the commit.

More notesEnter Game