Qwen2.5-7B-Instruct: a free hosted api for chat and coding agents
An OpenAI-compatible API for chat and coding agents on a free Kaggle GPU.
Problem it solvesPaid GPU hosting is too expensive for trying out a 7B model.
Omkar Chebale reports serving Qwen2.5-7B-Instruct on a free Kaggle T4 as an OpenAI-compatible HTTPS API for Postman, a backend, and a coding agent. He reports the setup handled tool calls, with 3,000+ requests used to find failure points.
How they grewKaggle free GPU and ngrok free tier
Omkar Chebale
@omkarchebale
I served a 7B LLM on a free GPU, made it work with a coding agent, and pushed 3,000+ requests through it to see where it breaks. $0 cloud bill. Here's everything I learned 👇 What I built Qwen2.5-7B-Instruct (4-bit AWQ) on vLLM, running on a free Kaggle T4. It's exposed as an OpenAI-compatible HTTPS API, so Postman, my backend and even a coding agent (with tool calls) plug straight in. The journey (everything that broke) 1/ The model wouldn't start. A T4 has no bf16 support, and Qwen defaults to bf16. One flag, --dtype half, fixed it. The 4-bit AWQ weights take ~5.5 GB of 16 GB, leaving room for the KV cache. 2/ The server kept dying. "Background" processes die when a Kaggle notebook finishes, and a pushed notebook finishes right away. I wrote a supervisor that restarts vLLM and the tunnel, and exits cleanly before Kaggle's 12-hour limit so the logs are saved. 3/ My secrets disappeared. Kaggle Secrets vanish silently when you push a notebook from the CLI. The fix: a private dataset holding the keys, created by one setup script. 4/ My coding agent's first request failed. "auto tool choice requires --enable-auto-tool-choice". vLLM needs an explicit tool-call parser, and a 4K context is useless for an agent. With the Hermes parser and 32K context, tool calls worked. 5/ I couldn't see anything. I put a small proxy in front of vLLM. It adds API-key auth (vLLM's /metrics endpoint had been public), a full trace per request, and GPU and KV-cache metrics every 15 seconds. 6/ My benchmarks lied to me. The first results: 13 seconds to first token. Awful. Then I found three benchmarks had been running at once, and one of them fired 512 requests together. In isolation, the same load gave 1.1 seconds. Never benchmark two things at once. The numbers (GuideLLM, through the public URL, network included) ⚡ 1 user: first token in 0.8s, then up to ~38 tokens/s. Faster than you can read. 👥 ~4 users at once: ~1.1s to first token, ~25 tokens/s each. 🚦 12+ at once (rough, from an early run): requests queue, and waits grow to 10 to 20s. 📈 Peak: ~215 output tokens/s across 16 parallel requests (measured directly on the server). In real terms: about 40 casual chat users, or 2 to 3 people running coding agents, on one free GPU. The surprise Under load the GPU showed 100%, and I assumed I'd hit the ceiling. Wrong. The KV cache used 2.5% of its reserved memory on average, and Kaggle's machine had a second T4 sitting completely idle the whole time. "100% GPU" told me one chip was busy, not that I'd used my hardware. Measure before you scale. Honest limits 12-hour sessions, 30 GPU hours a week, ngrok free-tier caps, and about 15 minutes to boot. A 7B model is great for chat and simple agent tasks, but not for complex multi-step coding. This is for learning, demos and small teams, not public traffic. Try it yourself It's open source. With a Kaggle account and a free ngrok account, one python setup.py file gives you your own endpoint in about 25 minutes:
Requests
3K+
through it



