A Codex harness that routes tasks and manages tool use
An alpha harness for Codex that routes tasks, controls tools, and reduces model choice for heavy users.
Problem it solvesHeavy GPT users keep hitting subscription limits and wasting time choosing models.
Teo Nex built a harness around Codex for heavy users. It routes tasks, manages tools and context, and preserves conversation history while aiming to reduce subscription-wall friction.
Teo Nex
@TeoNexcore
If OpenAI puts a bounty on me, this is probably why. The $200 plan was supposed to be enough. You start at $20. It covers your side projects and work tasks. Then you hit the limit, upgrade, and eventually you're paying $200 a month. Finally, enough room. Strongest model, maximum reasoning, more projects. Then you hit that limit too. “Holy shit, did I really get that much done?” In my case, the work hadn't changed enough to explain it. I've been paying for the $200 GPT subscription since April. My tasks and standards stayed roughly the same. How much I could get out of the subscription did not. April: I couldn't hit the cap. Didn't even know where it was. May: hit it for the first time. July: regularly running out toward the end of the week. September: spending far too much time deciding which model I could still afford to use. Yes, I'm a heavy user. I have 50–100 active sessions across a day. People see my workspaces and ask what the fuck is going on. But I wanted to keep working, so I built a harness around Codex. First: stop making me choose a model before every message. A lightweight router gets a short task description and picks the working model and reasoning level. The working model gets the full conversation. For a continuation, the harness can reuse the previous routing decision. Next: tools and context. In automatic mode, safe reads pass without another model call. Other actions go through a checking hook. Deleting or moving files, downloads and package installs are blocked in that mode. Manually selected models keep normal Codex permissions. Browser and desktop helpers return concise results instead of making the main model read huge interface dumps. Long tool outputs get shortened, with originals stored separately. For context compaction, the helper chooses a plan and Codex does the compression. Separate workers run their own conversations and return results and work logs. Background workers in the standard CLI are currently limited to tasks that don't write files. Then there's provider routing. My setup includes Codex, Google accounts, Opus, Gonka and Wally. I'm building quota-based fallback through OmniRoute while preserving conversation history. That part is still being finished; the distributable alpha starts with Codex. What did I actually pay? The $200 GPT subscription stayed. I also paid 8$ for eight Google accounts with 18 months of access, at 1$ each. Yes, from an unofficial seller. That's my purchase price, not Google's retail price. Gonka and Wally are free in my setup; Opus goes through Google. In this configuration, I wasn't paying an additional metered API bill. API-equivalent usage isn't money I actually spent. My actual expenses are the subscription and the account purchases above. This is still a working alpha, but it gets me much closer to covering my workload without constantly hitting the subscription wall. And I finally got to spend 5 billion tokens a week again I'm testing setup on other machines before releasing the code.





