shipitfaststupid

No name or maker matches “”. Press Enter to search descriptions too.

Catalog / Dev tools

0328X posts

Qwen2.5-7B-Instruct: a free hosted api for chat and coding agents

An OpenAI-compatible API for chat and coding agents on a free Kaggle GPU.

Problem it solvesPaid GPU hosting is too expensive for trying out a 7B model.

Omkar Chebale reports serving Qwen2.5-7B-Instruct on a free Kaggle T4 as an OpenAI-compatible HTTPS API for Postman, a backend, and a coding agent. He reports the setup handled tool calls, with 3,000+ requests used to find failure points.

How they grewKaggle free GPU and ngrok free tier

Omkar Chebale

@omkarchebale

I served a 7B LLM on a free GPU, made it work with a coding agent, and pushed 3,000+ requests through it to see where it breaks. $0 cloud bill. Here's everything I learned 👇 What I built Qwen2.5-7B-Instruct (4-bit AWQ) on vLLM, running on a free Kaggle T4. It's exposed as an OpenAI-compatible HTTPS API, so Postman, my backend and even a coding agent (with tool calls) plug straight in. The journey (everything that broke) 1/ The model wouldn't start. A T4 has no bf16 support, and Qwen defaults to bf16. One flag, --dtype half, fixed it. The 4-bit AWQ weights take ~5.5 GB of 16 GB, leaving room for the KV cache. 2/ The server kept dying. "Background" processes die when a Kaggle notebook finishes, and a pushed notebook finishes right away. I wrote a supervisor that restarts vLLM and the tunnel, and exits cleanly before Kaggle's 12-hour limit so the logs are saved. 3/ My secrets disappeared. Kaggle Secrets vanish silently when you push a notebook from the CLI. The fix: a private dataset holding the keys, created by one setup script. 4/ My coding agent's first request failed. "auto tool choice requires --enable-auto-tool-choice". vLLM needs an explicit tool-call parser, and a 4K context is useless for an agent. With the Hermes parser and 32K context, tool calls worked. 5/ I couldn't see anything. I put a small proxy in front of vLLM. It adds API-key auth (vLLM's /metrics endpoint had been public), a full trace per request, and GPU and KV-cache metrics every 15 seconds. 6/ My benchmarks lied to me. The first results: 13 seconds to first token. Awful. Then I found three benchmarks had been running at once, and one of them fired 512 requests together. In isolation, the same load gave 1.1 seconds. Never benchmark two things at once. The numbers (GuideLLM, through the public URL, network included) ⚡ 1 user: first token in 0.8s, then up to ~38 tokens/s. Faster than you can read. 👥 ~4 users at once: ~1.1s to first token, ~25 tokens/s each. 🚦 12+ at once (rough, from an early run): requests queue, and waits grow to 10 to 20s. 📈 Peak: ~215 output tokens/s across 16 parallel requests (measured directly on the server). In real terms: about 40 casual chat users, or 2 to 3 people running coding agents, on one free GPU. The surprise Under load the GPU showed 100%, and I assumed I'd hit the ceiling. Wrong. The KV cache used 2.5% of its reserved memory on average, and Kaggle's machine had a second T4 sitting completely idle the whole time. "100% GPU" told me one chip was busy, not that I'd used my hardware. Measure before you scale. Honest limits 12-hour sessions, 30 GPU hours a week, ngrok free-tier caps, and about 15 minutes to boot. A 7B model is great for chat and simple agent tasks, but not for complex multi-step coding. This is for learning, demos and small teams, not public traffic. Try it yourself It's open source. With a Kaggle account and a free ngrok account, one python setup.py file gives you your own endpoint in about 25 minutes:

Requests

3K+

through it

Sep 30, 2026Read on x.com ↗

Also filed under Dev tools

  1. 0426

    tiun: one system for auth, payments, customer data and analytics

    The main mission behind @tiun_io is simple: Help builders focus on building better products, not managing more tools. Before we started building tiun, we spoke with 150 founders. Some conversations happened online. Others happened in person. That gave us a huge advantage. We could see where founders were losing time, what they were connecting manually, and what kept breaking. That is why tiun is built as one system for auth, payments, customer database, and analytics. The goal is to replace several tools in your stack. Less time connecting infrastructure. More time focused on the actual product. But building the product is only one part. You also need users. That is one reason we started running hackathons. Our first one happened in August together with @SaidAitmbarek . More than 800 builders joined. More than 120 products were submitted. Yesterday, we announced the winners. The final result combined community voting with a jury reviewing the ideas and products. The goal was not just to pick winners. We wanted to give builders more visibility and help their products grow beyond the hackathon itself. Then, 9 days ago, we launched tiun on Product Hunt. We spent around two months preparing for that launch. It ended with tiun becoming Product of the Day. That launch brought another wave of builders to tiun. More users means more feedback. More feedback means more ideas for what we should build next. And it gives us even more motivation to make tiun valuable for builders. So I’m really happy to keep building for you. You keep growing the products you are building. We will keep growing the product that helps you build them. tiun.io/?utm_source=x&…

    @NikChr1414 · Dev tools

  2. 0421

    A Chrome extension with a pitch that missed the mark

    I posted my Chrome extension on r/ClaudeAI. 4,000 views 9 clicks 0 installs 0 upvotes What I got wrong: → price in the post → feature list, no story → paid tool, pitched to devs who build their own Next post leads with a free tool. I'll share the numbers.

    @mnagm93 · Dev tools

  3. 0412

    A SaaS making a small monthly income

    My SaaS finally making me $10 a month 🤑 #buildinpublic

    @i0x46 · Dev tools

  4. 0406

    A subscription app being scaled to recurring revenue

    day 30 of scaling my app to 10k MRR $17 away from $100 MRR. 22 active subscriptions. $713.70 revenue over the last 28 days. Still early in the $10k/month journey. Seeing people pay for something I built feels good. Sharing these small steps is part of why I started building in public :)

    @Artur_Abra · Dev tools

  5. 0405

    A developer tool with transaction tracking

    This month has been a quiet one. One my app did Over 4,000 transactions came through, but revenue is still sitting around ₦772k. Not the kind of month I was hoping for, but I’ve learned not to judge the whole journey by one slow month. We keep pushing.

    @lhilwick · Dev tools

  6. 0402

    Tollbooth DPYC: monetization service for MCP tool calls

    ☕ At Good Brew, we bring to market high-quality coffees, merchandise for coffee preparation, and books that spread the Gospel of Sound Money — to a world desperate for salvation from the sins of Keynesian central bankers. ₿ Bitcoin advocates know that it is wise to hold onto the BTC they have acquired. We stood up Good Brew — not only to share our love of coffees — but to help entrepreneurs learn how to honor Gresham's Law, to trade away bad money to acquire good money. That is, through the Tollbooth DPYC ecosystem, we help entrepreneurs turn the new digital services they have conceived into value that can be easily and quickly exchanged for BTC — not only to HODL bitcoin, but to earn it. 📚 At our Shopify store, we publish a variety of free articles on our opinions on coffees and, more importantly, our experience with the technology of modern digital finance. Here is a summary from one of the articles. Please come read more and explore the site. Cold-start latency is residual friction on the path between an autonomous agent and a Bitcoin-settled Operator. This piece reports the work of moving a Tollbooth-DPYC Operator onto Akamai's Spin WebAssembly edge so that cold start falls to roughly 0.4 ms. The protocol library is left untouched. Only the thin server shell changes. Three gaps between CPython and the Wasm edge — TLS termination, elliptic-curve crypto, and a Nostr WebSocket relay — are closed with stand-ins that hold no secrets. The Operator's sole private key stays its own; credentials arrive encrypted and Authority-signed at runtime. A weather Operator already proves the pattern. Sound money favors low time preference. The infrastructure under it should favor low latency. Read more at cafe.tollbooth-dpyc.com/blogs/our-valu… ──────────────────── ₿ Tollbooth DPYC™ is the efficient monetization service for MCPs whose BTC-using entrepreneurs want to avoid KYC and who agree to practice Don't Pester Your Customer™ commerce. ⚡ With Tollbooth DPYC, your patrons purchase, out of band, tranches of "𝐴𝑃𝐼 𝑆𝑎𝑡𝑠", using rapid-settlement Lightning. Patrons then spend these API Sats across different tool calls with zero per-request friction and no KYC. 🏛️ Accounting of available balances is via voluntary, arbitrary, patron-chosen Nostr-signed npubs, without PII, KYC, or KYT. As Operator, you collect your revenue in your wallet, without any KYC with a traditional bank. ◆ Open source, Apache 2.0 — github.com/lonniev/tollbo… ◆ Patent pending (US Prov. App. No. 64/045,999) ◆ Pricing Studio for operators → tollbooth-dpyc.com Posted by Excalibur-MCP, a Tollbooth DPYC™ MCP service.

    @lonniev · Dev tools

  7. 0392

    A merchant of record for indie app payments in India

    wanted stripe for my app. india said no, not without a GSTIN, and im nowhere near ₹20 lakh revenue. ended up with a merchant of record instead. they sell it, handle the tax, pay me out. nobody tells indian solo devs this on day 1 @stripe @dodopayments

    @yadavpiyush005 · Dev tools

  8. 0382

    A e-signature app that got abused as a phishing relay

    A spammer turned my e-signature app into a free phishing machine🤯. They used my domain to send 1000s of phishing emails. Found out when Resend emailed me that my free quota was used up. Mistakes I made: - No limit on recipients per document (one doc went to 664 people) - No canonical email check. [email protected] goes to the same Gmail inbox as [email protected], but my app treated it as a new user with a fresh free quota Fixes: - Suspended all the accounts - Capped recipients per document - Accounts that share one inbox now share one quota - Blocked the signing links the victims already got Beware: If your app sends emails on behalf of users, you're a spam relay until you prove otherwise.

    @heyakbarali · Dev tools