Is OpenCode good? We ran it on real bugs for $0
Key takeaways
- In our September 15, 2026 test, OpenCode v1.18.31 fixed two real bugs in the ky HTTP client using only free models, with no account and no API key. Big Pickle passed the maintainer’s tests on both; MiMo-V2.5 Free passed on one of two.
- MiMo-V2.5 Free’s failed fix passed the tests it wrote for itself but not the maintainer’s, and it didn’t compile under the library’s build config. It never ran the type checker on that task.
- OpenCode’s free models work without signing in, although its Zen docs describe adding billing details first. You pay with data instead: the docs say the free models may use what you send “to improve the model.”
- Out of the box OpenCode edits files and runs shell commands without asking. Big Pickle ran 25 shell commands across the two tasks, including
npm run buildandgit stash, without a single prompt. - Stock OpenCode loaded 8,068 tokens of context on the first step of a one-line prompt. On a machine with Claude Code skills installed that rose to 13,854, and one permission rule brought it back down to 7,746.
OpenCode is the open-source (MIT) coding agent that runs in your terminal and talks to whichever model provider you point it at. It’s good, and its free tier is better than we expected, with one catch worth planning around. Three of four free-model runs produced fixes the maintainer’s own tests accepted. The fourth fixed the symptom in the bug report, wrote tests that agreed with it, and left code the project’s build rejects. The weak spot is verification: the agent checks less than you’d assume, and runs whatever it likes on your machine while it does.
How we tested OpenCode
We gave OpenCode real bugs whose fixes it couldn’t see, and graded the result with tests it never saw either. ky is a small TypeScript HTTP client with a thorough test suite, and two of its bugs were fixed at the end of August 2026:
- Task 1: issue #879.
replaceOptiondoesn’t replace anAbortSignalwhen you extend an instance, so aborting the parent still aborts the child. Fixed in PR #880. - Task 2: issue #878.
onDownloadProgressnever reports 100% for an empty response body. Fixed in PR #881.
For each task we checked out the commit before the fix and handed the agent the user’s bug report and repro. Both reporters had also written up the cause and a proposed fix, and we cut those, because leaving them in turns the task into typing. The prompt asked for a fix plus tests, told the agent to run the relevant tests to confirm they pass, and gave it the command for a single test file. It said nothing about type checking or linting.
Grading was the tests the maintainer merged with the real fix, copied in afterwards from outside the agent’s directory, plus a type check of the library source. We also ran ky’s linter and report what it found, but it didn’t decide pass or fail. The grader was checked both ways before any agent ran: the new tests fail on the old code (3 failures on task 1, 2 on task 2) and pass once the official fix is applied.
We installed OpenCode v1.18.31, released September 14, with brew install anomalyco/tap/opencode on an Apple M2 Mac with 16 GB of RAM. The binary is 137.8 MB, and opencode --version took 1.73 seconds cold and 0.28 seconds after that. Each task ran headless through opencode run --format json with two free models, MiMo-V2.5 Free and Big Pickle. As a reference point, Claude Code on Opus 5 ran the same prompts through claude -p, with an allowlist we wrote (file tools, plus git, npx, node and read-only shell commands) so it never stopped to ask either. Token counts below are what each model’s provider reported per step, so they use that model’s own tokenizer.
Did the free models fix the bugs?
Three of four runs did. Big Pickle passed everything on both tasks. MiMo-V2.5 Free passed task 1, including a test that checks a signal key nested inside a JSON body is left alone, and failed task 2.
| Lane | Task 1 (#879) | Task 2 (#878) | Cost to us |
|---|---|---|---|
| OpenCode + MiMo-V2.5 Free | Passed · 108 s · 18 steps | Failed · 317 s · 13 steps | $0, no account |
| OpenCode + Big Pickle | Passed · 264 s · 36 steps | Passed · 158 s · 30 steps | $0, no account |
| Claude Code + Opus 5 | Passed · 202 s · 20 turns | Passed · 161 s · 21 turns | Subscription ($1.08 and $0.90 at API rates) |
The miss taught us the most. On task 2, MiMo changed ky’s progress stream so the final 100% event always fires, which is exactly what the bug report asked for. For an empty body, though, it passed undefined as the chunk, and ky’s callback type promises a Uint8Array. The maintainer’s tests check the chunk. MiMo’s own tests checked the percentage and the byte counts, so they passed. It did what the prompt asked: it ran that test file, saw green, and stopped. ky’s tests execute through tsx, which strips types without checking them, so nothing it ran could catch the mistake. Under the library’s tsconfig.dist.json the change throws two TS2345 errors, and the tsc step of npm run build exits with code 1.

We nearly scored it wrong ourselves. Our first type-check pass used ky’s tsconfig.test.json. That config only covers the type tests in test-d/ and never looks at the source. Claude Code made the same choice on both tasks, and its fixes happened to compile. Big Pickle made it on task 1 as well, but it had already run npm run build, whose tsc step does check the source, so it was covered by accident. We weren’t, until we went back.
On task 1 the passing fixes were close to the maintainer’s. MiMo and Claude Code both unwrap replaceOption before the signal check and clear inherited signals on a replacement, undefined or null. That’s the official fix’s shape without its root-level guard, plus null handling the bug report asked for and the official fix leaves out. We went looking for a regression from the missing guard, a signal: null nested in a JSON body, and found none: the request body came out identical on the unfixed code, the official fix, and all three agents’ versions. Big Pickle’s fix was twice the size of the maintainer’s (28 added lines against 14), added a TypeError nobody requested, and left five new lint warnings, against three for MiMo and none for Claude Code.
Two tasks can’t rank models. Both bugs were small. Wall times are indicative, since some lanes ran at the same time and free-tier latency swings: 117 of MiMo’s 317 seconds on task 2 went on waiting for its first response.
Is OpenCode really free?
The agent is free, and the free models don’t need an account, but they come with conditions. Before we logged in to anything, opencode models listed seven free models, and opencode run -m opencode/mimo-v2.5-free answered a one-word prompt in about 12 seconds. The Zen page describes a gate that wasn’t there: “You sign in to OpenCode Zen, add your billing details, and copy your API key.” The code explains it. When we read OpenCode’s source, the registration step with no credentials deleted every model whose input cost isn’t zero. So a fresh install gets exactly the free models, and our run confirms that’s what ships.
What you pay with is data. The same page says of the free models that “During its free period, collected data may be used to improve the model,” and it labels the two NVIDIA endpoints “Trial use only — do not submit personal or confidential data.” The starkest condition belongs to the Muse Spark contributor model: “Heavily discounted token pricing in exchange for permission to use your prompts and completions to train future Meta models.” Two Muse Spark contributor models showed up in our free list. Each free model on the Zen page is available “for a limited time,” with no end date and no published usage limit. We ran them on an open-source repo on purpose. We wouldn’t run them on a client’s code.
The other free route is bringing a key or a subscription you already pay for. The agent is MIT-licensed and costs nothing either way; you pay your provider directly.
Can you use your ChatGPT, Copilot or Claude subscription?
ChatGPT and GitHub Copilot, yes. Claude Pro and Max, no. The providers page lists “ChatGPT Plus,” “Github Copilot” and “Gitlab Duo” as subscriptions that work “with zero setup,” with a browser login for ChatGPT Plus/Pro and a device-code login for Copilot. It adds that some Copilot models “might need a Pro+ subscription.”
On Claude it’s blunt: “There are plugins that allow you to use your Claude Pro/Max models with OpenCode. Anthropic explicitly prohibits this.” OpenCode stopped bundling those plugins as of version 1.3.0. If Claude is the model you want, the sanctioned routes are an Anthropic API key or Claude Code itself. The pairing also works the other way round: the Go docs list Claude Code as a validated client for OpenCode’s Go models.
We didn’t test the ChatGPT or Copilot logins, or Claude Code against Go. Those paragraphs are the docs, not our run.
Is OpenCode Go worth it?
For bugs this size, you wouldn’t be buying quota. Pay for Go if you want its stronger models or want your prompts kept out of training. OpenCode Go, the flat subscription in the pricing box above, sets its limits per model in dollars of usage: “Each model has the following usage limits: 5-hour — 20% of the monthly limit; weekly — 50%; and monthly — 100%.” As of September 15 the monthly limits are $15, $30 or $60 depending on the model. That structure is less than a week old, and the docs warn that “Usage limits may change as we learn from early usage and feedback.”
Here’s the arithmetic, and it is arithmetic, not a measurement. MiMo-V2.5 Free’s passing run on task 1 used 30,855 input tokens, 4,556 output tokens and 439,744 cached-read tokens. Go lists MiMo V2.5 at $0.14 input, $0.28 output and $0.0028 cached read per million tokens, with a $60 monthly limit, or $12 every five hours. That prices the fix at about $0.0068, so roughly 1,750 fixes of that size fit under the five-hour limit. Go’s limits are in dollars, not requests; its own estimate of about 30,100 requests per five hours for that model works out to around 1,670 runs of 18 steps. The free model and Go’s MiMo V2.5 aren’t guaranteed to be the same deployment, so treat it as an order of magnitude.
The case for Go we couldn’t test is reliability. The docs pitch it as “reliable access to popular open coding models,” and our free-tier runs had one 117-second wait for a first response. We didn’t subscribe, so whether Go avoids that is still an open question for us.
Per-model limits are also the part people argue about. On September 14 a Reddit user said they’d subscribed expecting $60 of DeepSeek V4.1 Flash usage and found $15 in the docs. The docs table now shows both figures on that row: $15 struck through, and $60 marked “4x · Ends Sep 20.” That promotion row reached the docs on September 13. One commenter asked for more transparency, saying limit changes tend to surface only on X; we haven’t checked where past changes were announced.
Is OpenCode safe to use on your machine?
Only once you’ve configured it, because the defaults trust the agent with nearly everything. The permissions page says “If you don’t specify anything, OpenCode starts from permissive defaults,” with most actions set to allow. Touching paths outside the working directory asks first, and so does repeating an identical tool call three times. Reading .env files is denied. Editing files and running shell commands aren’t gated at all.

Our logs show what that looks like. Across the two tasks Big Pickle ran 25 shell commands without a prompt, among them npm run build, which deletes and regenerates ky’s distribution folder, and a git stash followed by git stash pop to compare lint output against the untouched code. Nothing broke. It was sensible engineering by a stealth model we’d never used before, rewriting git state in our checkout without asking. Here’s the counter-case: Claude Code stashed too, on both tasks, but only because our allowlist pre-approved git. OpenCode needed no allowlist to get there.
OpenCode’s provider also brings a web search tool you didn’t turn on. The tools page says websearch is available “when using the OpenCode or OpenCode Go provider,” so it comes with the free models and the paid Zen ones alike. Big Pickle used it once on task 2, searching for “ky HTTP client onDownloadProgress empty body percent 1 flush fix github.” That query describes your code, and it leaves your machine for a search provider. In a test like ours it’s also a leak risk, so we read all nine results. None contained the issue or the fix.
The zero-config safeguard is Plan mode. The docs describe it as a mode “that disables its ability to make changes and instead suggest how it’ll implement the feature,” and Tab toggles it in the terminal UI. We ran headless, so we didn’t use it. For real work, a few lines of opencode.json keep file edits and shell commands behind a prompt while letting routine commands through:
{
"$schema": "https://opencode.ai/config.json",
"permission": {
"edit": "ask",
"bash": {
"*": "ask",
"npx ava *": "allow",
"git status*": "allow"
}
}
}
The last matching rule wins, which is why the catch-all "*" goes first. For what OpenCode’s own services see and which headers ride along, see our source read of v1.18.28. The short version from that code: with your own API key, OpenCode’s session headers never attach.
How many tokens does OpenCode load before your prompt?
About 8,000 on a clean install, and more if you already use Claude Code. With HOME pointed at an empty directory and a one-line prompt, MiMo-V2.5 Free reported 8,068 tokens on the first step, identical across two runs. It’s the first OpenCode number we’ve measured ourselves. When we measured Claude Code’s startup overhead, we could only cite other people’s figures for OpenCode.
From a normal home directory the same prompt came in at 13,854. The skills docs explain it: “Global definitions are also loaded from ~/.config/opencode/skills/*/SKILL.md, ~/.claude/skills/*/SKILL.md, and ~/.agents/skills/*/SKILL.md.” Our test machine has 44 distinct skills: 31 in ~/.agents/skills, and 13 more in ~/.claude/skills, where the other 31 entries are symlinks back to the agents directory. OpenCode listed every one of them in its skill tool’s description without a word. The skill bodies load only on demand; the catalog rides along with every request. OPENCODE_DISABLE_CLAUDE_CODE=1 only got the count to 12,021. That variable covers .claude, and no environment variable covers .agents.
A permission rule did the job. With "permission": {"skill": {"*": "deny"}} in opencode.json, the same prompt on the same loaded machine used 7,746 tokens, a little under the clean install. We haven’t pinned down why it’s lower; our guess is that denying every skill also drops the skill tool’s own definition. Vercel’s fx borrows skills the same way, so this is a convention spreading across agents. For us it was still about 5,800 tokens per session that nobody opted into.
Is OpenCode as good as Claude Code?
Close, on free models, and Claude Code still came out ahead. Claude Code on Opus 5 passed both tasks. OpenCode on free models passed three of four runs. What separated them was verification: on the task MiMo failed, it ran the test file the prompt pointed to and stopped. Claude Code read the commit that introduced the progress code, ran every test file that doesn’t need a browser, ran the linter, and stashed its change to compare type-check output against the untouched code before calling it done. Big Pickle checked nearly as much, and passed.
This design can’t split the harness from the model. MiMo’s thin checking on task 2 belongs to that run of that model inside OpenCode, and the same model checked more on task 1, where it ran the type checker unprompted. Our Claude Code lane also ran on a working setup with its own skills, plugins and allowlist, so it isn’t a stock comparison. What holds is narrower: if you’ve been assuming OpenCode’s free tier is too weak for real work, on bugs this size it isn’t. Just don’t take a green test run it wrote for itself as the finish line.
How to reproduce our run
Everything here is public. Clone ky, check out 294fe63 for task 1 or 0bda554 for task 2, and install with npm install --ignore-scripts. Give the agent the problem and repro sections of the issue, followed by: “Fix the bug in the library source, add tests that cover it to the existing test suite, and run the relevant tests to confirm they pass,” and the command for a single test file (npx ava test/<name>.ts). Afterwards, copy test/main.ts (task 1) or test/stream.ts (task 2) from the fix commit into the working tree and run npx ava on that file, then npx tsc --noEmit -p tsconfig.dist.json for the source type check. Token counts come from the step_finish events that opencode run --format json prints.
For the context numbers, run opencode run -m opencode/mimo-v2.5-free --format json "Reply with exactly one word: pong" three ways: with HOME set to an empty directory, from your normal home, and with the skill deny rule in a project opencode.json. Add tokens.input and tokens.cache.read from the first step_finish event.
More hands-on testing of coding agents lives in our reviews.
Frequently asked questions
What is the limit for free models on OpenCode?
OpenCode's docs at v1.18.31 don't publish one. The Zen page lists each free model at a price of Free and says it's available "for a limited time," with no request or token cap, unlike the per-model dollar limits it publishes for the paid Go plan. The free models finished every run in our test, though MiMo-V2.5 Free once waited 117 seconds for its first response.
Can OpenCode replace Claude Code?
On our two bugs it came close on free models: three of four OpenCode runs produced fixes the maintainer's tests accepted, and Claude Code on Opus 5 passed both. That's two tasks, not a benchmark. You'd give up Claude itself, since OpenCode's docs say Anthropic prohibits using a Claude Pro or Max subscription inside OpenCode, and you'd want a permission config, because OpenCode edits files and runs shell commands without asking by default.