Max out free quotas first, then decide whether to pay—but a large number of online "free-tier" tutorials are already outdated, and following them will waste your time. From text model APIs, coding agents, and image/video generation, to search crawlers, code sandboxes, memory layers, and free cloud deployment, this is an operating manual still valid as of August 2026 with 137 primary source links. Every entry follows the format: "open link → do what → get how much → what's the limit."
Before writing this, I specifically verified the "classic free-tier paths" circulating online, and found that a significant portion have already expired, or look free but are no longer what they seem. If you're about to follow an old tutorial, check this table first:
| Channel | What you could previously get | Current status |
|---|---|---|
| Gemini CLI | Personal Google account login, thousands of free requests/day | Service stopped for all accounts on June 18, 2026; migrated entirely to closed-source Antigravity CLI ↗ |
| GitHub Models | Free model calls via GitHub account | New user access stopped June 16, 2026; fully retired for all users July 30 ↗ |
| OpenAI Sora | Free quota for video generation | Free quota canceled January 2026; web/App shut down April 26; API to terminate September 24 ↗ |
| Yi API (01.AI) | Register to call Yi series models | New user registration stopped August 10, 2026; full shutdown September 3 ↗ |
| PlayHT | Free voice cloning/TTS | Acquired by Meta July 2025 ↗; consumer product closed end of that year |
| Bing Web Search API | Free quota for Bing search to connect agents to the web | Fully retired August 11, 2025 ↗; no new registrations, all old keys invalid |
| Brave Search API | 2,000-5,000 free searches/month | Fixed free tier canceled February 2026 ↗; replaced by auto-credited $5/month (≈1,000 searches), overage incurs charges |
| Google Programmable Search Engine | 100 free searches/day | Registration closed to new users January 2026 ↗; full shutdown January 1, 2027 |
| Railway | Free cloud deployment quota on registration | Now only a 30-day Trial with one-time $5 credit; Free plan is $1/month ↗; real production use basically requires payment |
| Fly.io | 3 shared-CPU VMs free long-term | Official statement: "there is no free tier"; now only 2 VM-hours or 7-day Free Trial ↗ |
| Cursor student free year | .edu email verification grants one year of Pro | Closed to new users June 25, 2026 ↗ (officially citing massive fraud abuse); existing holders retain until expiry |
| GitHub Copilot student unlimited | Unlimited Copilot Pro after student verification | Changed to 200 credits/month March 2026 ↗; new applications suspended April 20 |
Below, channels that are still alive are listed by category. Every entry follows the format: "open link → do what → get how much → what's the limit."
OpenRouter—open openrouter.ai ↗, filter for models with the :free suffix (currently ~17 on offer, mainly Gemma, GLM, MiniMax, Nemotron, etc.; list scrolls with updates ↗, refer to the real-time list). Free tier rate-limited to 20 req/min; the key operation is to top up your account with at least $10 (even just once), and the daily cap permanently jumps from 50 to 1,000—this is a lifetime unlock, not a balance threshold (official rate-limit docs ↗, rules explanation ↗).
Occasionally, time-limited free top-tier "stealth models" pop up—the latest is Ox Alpha ↗, launched August 20, 2026: 1M-token context, purpose-built for coding and agentic tasks, OpenCode gave it nearly unlimited free usage for a week ↗. Historically, Quasar Alpha (→GPT-4.1) ↗, Horizon Alpha (→GPT-5) ↗, and Sonoma Alpha (→Grok 4 Fast) ↗ were all unmasked within weeks and either went paid or delisted—use them if you spot them, but don't treat them as long-term channels. This example played out especially fast: on August 26, Zhipu (Z.ai) confirmed to media that Ox Alpha was actually their new flagship GLM-5.3-Flash in disguise ↗; the free stealth route was promptly delisted from OpenRouter and replaced with a paid official API.
NVIDIA NIM—register at build.nvidia.com ↗, no credit card required; 1,000 credits auto-credited, enterprise email verification can request 5,000; rate ~40 RPM, covering 80+ models including Llama 4, Qwen, Hermes 3 (developer forum credit discussion ↗).
Groq Cloud—register at console.groq.com ↗; free tier runs open-weight models like gpt-oss-120b/gpt-oss-20b/qwen3.6-27b, rate-limited to 30 RPM / 1K RPD / 8K TPM / 200K TPD; adding a credit card (no charge) boosts limits 10×.
Google AI Studio (Gemini API)—ai.google.dev ↗ register to get free quota for Gemini Flash/Flash-Lite series; however, since April 2026 the Pro series has been removed from the free tier, and rate limits no longer publish fixed numbers—check the console for actual quotas.
Mistral La Plateforme (Experiment tier)—mistral.ai/pricing ↗ free tier covers all models (including Mistral Large, Codestral), limited to non-production evaluation; check Admin Console for specific rate limits.
These providers share a common trait: none require enterprise credentials; phone number / personal identity verification gets you basic free quota, and enterprise verification only raises the quota ceiling.
| Provider / Product | Free Content | Threshold | Source |
|---|---|---|---|
| Zhipu GLM-4.7-Flash | Permanently free, 30 concurrent cap, no expiry date | Personal registration | Pricing page ↗ |
| Baidu Qianfan / ERNIE ERNIE Speed/Lite/Tiny | Pay-as-you-go free long-term since May 2024 (non-flagship models) | Personal or enterprise identity verification | Free opening announcement ↗ |
| iFlytek Spark Lite | Officially permanently free; plus 2M tokens new-user quota pack (Pro/Max) | Phone number registration | iFlytek open platform ↗ |
| Tencent Hunyuan | 1M tokens / 12 months after identity verification; TokenHub new-user promo through Dec 31, 2026 | Tencent Cloud personal identity verification | Billing docs ↗ |
| ByteDance Doubao / Volcengine Ark | "Worry-free trial" gifts 500K tokens per model; face verification adds another 500K / 30 days | Personal identity (face verification) | Activation management docs ↗ |
| Alibaba Qwen (Bailian) | New users ~70M tokens (1M each across 70+ models, valid 90-180 days) | Alibaba Cloud identity verification | First-call instructions ↗ |
| DeepSeek | New account ~5M tokens (30-day validity; exact number not statically confirmed on official page) | Phone / email registration | Pricing info ↗ |
| Kimi (Moonshot) | Registration grants ¥15 trial credit (~15M tokens) | Phone number registration | Pricing docs ↗ |
Don't confuse them: Zhipu's GLM-5.3-Flash, released August 26, 2026, is the new flagship and is not free—320B-parameter native multimodal model, MIT-licensed open weights but API billed per token (50% off promo through Sept 9); the one you can free-ride in the table is GLM-4.7-Flash ↗—they are not the same thing.
MiniMax: No official first-party announcement found for "register and get a fixed token pack"; phone/email/WeChat registration ↗ works, developer门槛 is low, but quota info should be based on actual console display. ModelBest MiniCPM: Focused on open-weight model local/edge deployment ↗, not a stable commercial API free-tier channel; suited for those with local compute.
Hermes Agent (Nous Research, MIT-licensed open source)—hermes-agent.nousresearch.com ↗ deployed on your own server, the software itself is permanently free. It natively supports multiple model providers; the practical approach is to point the model backend at the free channels from sections 1 and 2—OpenRouter's :free pool, NVIDIA NIM's registration quota, GLM-4.7-Flash's permanently free key—so neither the software nor the model calls cost anything (free provider list ↗). The same logic applies to "bring-your-own-key" coding agents like Cline and Continue.dev: the tool itself is free, the model backend it connects to is free, so the entire agent is free.
GitHub Copilot Free—activate directly with a GitHub account; official quota ↗ is 2,000 code completions + 50 chat requests/month; model auto-assigned, cannot manually select.
Cursor Free (Hobby)—register and use directly; official pricing page only says "limited Agent requests / limited Tab completions," no longer publishes specific numbers; new accounts include a 1-week Pro trial. Windsurf Free—unlimited Tab auto-completions, Cascade agent quota roughly enough for 2-3 days of typical coding.
Two pitfall warnings: Gemini CLI is dead (see §00), and its successor Antigravity's free tier ↗ is itself shrinking fast—250 req/day at launch in November 2025, cut to 20 req/day in December 2025, and since March 2026 overage requires purchasing extra credits at $0.01/credit. Amazon Q Developer ↗ technically still works (50 agentic requests/month), but new user registration stopped May 15, 2026 and support ends April 2027—not worth learning now. ByteDance Trae's domestic version (trae.cn) is completely free with no门槛, with built-in Doubao and DeepSeek R1/V3; the overseas version has tiered limits, with reports of sudden tightening in July 2026—test it yourself before relying on it.
Building your own agent isn't just about calling models—you also need to give it "eyes and hands" (web search, web scraping), an "execution environment" (code sandbox), and "memory" (memory layer + vector database). These infrastructure components also have free tiers, and many tutorials haven't mentioned them at all.
Search & Web Scraping APIs
| Product | Free Quota | Notes |
|---|---|---|
| Tavily ↗ | 1,000 free credits/month | Search API designed for AI agents, no credit card required |
| Firecrawl ↗ | 1,000 free credits/month (1 credit = 1 page) | Official wording: "No cost, no card, no hassle" |
| Exa ↗ | $20 on registration + $10 auto-credited/month | No payment method needed to start |
| Serper.dev ↗ | One-time 2,500 queries (not monthly refresh) | No credit card required |
| SerpApi ↗ | 250 searches/month, 50/hour throughput cap | — |
| Jina AI Reader ↗ (r.jina.ai) | Completely free, no API key needed, 20 RPM; with key: 500 RPM + 10M tokens | Append URL directly after r.jina.ai/ to use |
| ScrapingBee ↗ | One-time 1,000 free API credits | No credit card required |
| You.com API ↗ | Keyless Web Search 100 free calls/day; developer account grants $100 credit | — |
| ScrapingAnt ↗ | 10,000 free credits/month, recurring long-term | No credit card required |
| Apify ↗ | $5/month platform credit | No credit card required; also application-based $500/6-month new-user promo |
Perplexity API (Sonar series) has no free tier—official pricing docs ↗ are entirely per-token/per-request billing with no mention of any free quota; new accounts must bind a payment method before generating a usable API key. The rumored "Pro subscription grants credit" claim is unconfirmed by official docs.
Code Execution Sandboxes
E2B ↗: Hobby tier one-time $100 usage credit (not monthly reset), max 1 hour/session, up to 20 concurrent sandboxes, 10GiB free storage, no credit card. Daytona ↗: one-time $200 free compute credit, no credit card. Modal ↗: $30 free compute credit/month, monthly reset non-cumulative (see §09).
Browser Automation Services
| Product | Free Quota | Notes |
|---|---|---|
| Browserbase ↗ | 3 concurrent browsers, 1 browser-hour/month, 15 min/session | Also includes 1,000 Search/Fetch API calls + $5 Model Gateway credit, no credit card |
| Steel.dev ↗ | One-time $30 credit (90-day validity, ≈300 browser-hours) | Free tier renamed "Launch"; many tutorials still cite the old "$10/month recurring credit"—outdated |
| Browserless.io ↗ | 1,000 units/month (≈8.3 browser-hours), 2 concurrent | No credit card required |
| Hyperbrowser ↗ | 5,000 credits, 1 concurrent, 7-day data retention | Official wording: "No credit card required" |
| Anchor Browser ↗ | 5 credits/month (≈50 tasks), max 5 concurrent | — |
| Browser-Use Cloud ↗ | $0/month, 3 concurrent, 10 agent tasks/month | Open-source self-hosted version completely free (MIT license ↗), just bring your own LLM API Key |
| AgentQL ↗ | 50 free API calls/month + 10 hours remote browser quota | No credit card required |
MultiOn/Multion ↗ has pivoted to consumer personal assistant product "AGI-0" and no longer offers a free tier for developer browser automation APIs; if tutorials still recommend it as an agent browser service, the info is outdated.
Vector Databases & Agent Memory Layers
General vector databases store and retrieve embeddings, while "memory layer" services like Mem0 and Zep additionally handle agent-specific logic such as memory extraction, summarization, and forgetting—different roles:
| Product | Type | Free Quota |
|---|---|---|
| Pinecone ↗ | Vector database | 2GB storage, 2M Write Units/month, 1M Read Units/month, no credit card |
| Qdrant Cloud ↗ | Vector database | 0.5 vCPU / 1GB RAM / 4GB disk, permanently free, no credit card |
| Weaviate Cloud ↗ | Vector database | 100,000 objects, 1GB memory, 10GB disk; free tier changed from old "14-day Sandbox" to permanently free "Engram" |
| Supabase ↗ | Database (with pgvector) | 500MB database, 5GB egress; projects auto-pause after 1 week idle, max 2 active projects |
| Turso ↗ | SQLite database | 100 databases, 5GB storage, 500M row reads/month, no credit card |
| Mem0 ↗ | Agent memory layer | Permanently free, 10K memory writes + 1K retrievals/month; official wording: "permanently free" |
| Zep ↗ | Agent memory layer | 10,000 credits/month (1 credit = 350 bytes); open-source self-hosted version maintenance stopped April 2025, only commercial Cloud remains |
| Letta (formerly MemGPT) ↗ | Agent memory layer | $0/month, limited to 3 stateful agents, supports BYOK unlimited calls |
| Supermemory ↗ | Agent memory layer | $5-equivalent credits/month; official wording: "No credit card required" |
| Cognee ↗ | Agent memory layer | 1 workspace + 1M tokens, "Free forever"; open-source version self-hostable permanently free |
Chroma Cloud ↗ grants a one-time $5 credit, but the official FAQ notes the Starter tier "requires a credit card"—a rare exception among these options. LangSmith ↗ (LangChain observability platform) free tier includes 5,000 traces/month but no free LangGraph Platform hosted deployment—if you want to free-ride "hosted agent runtime with memory," it can't help; only useful for viewing traces.
Colab/Kaggle in §09 are for "running experiments"; this section is about writing your agent as a long-running online service accessible to others—two different scenarios that people often confuse.
| Platform | Free Quota | Notes |
|---|---|---|
| Cloudflare Workers ↗ | 100K requests/day, permanently free, 10ms CPU time/invocation | No credit card, unchanged long-term, currently the most stable free deployment channel |
| Vercel (Hobby) ↗ | 100GB bandwidth, 1M requests/month, 4 CPU-hours | Official terms explicitly prohibit commercial use, limited to personal non-commercial projects; violations result in takedown without notice |
| Netlify ↗ | Changed to credit system September 2025, 300 credits/month | Overage pauses until next month, no bill generated, no credit card |
| Render ↗ | 512MB RAM Web Service, 750 free instance-hours/month per workspace | 15-min idle = sleep, 30-60s cold start; officially no credit card needed, but community reports frequently being asked to verify card |
| Deno Deploy ↗ | ~1M requests/month, 100GB bandwidth, 15 CPU-hours | Commercial use allowed, but card-verified organization required to use full quota |
| Supabase Edge Functions ↗ | 500K invocations/month | Shares account quota with database, not independent allocation |
Railway and Fly.io—see §00—these two formerly classic free deployment channels now exist in name only; don't follow old tutorials to free-ride them.
| Product | Free Content | Notes |
|---|---|---|
| Stability AI ↗ | 25 credits on new registration (~3 Stable Image Core images), no credit card | Pay after use-up, no monthly free quota |
| HuggingFace Spaces ↗ | Official Stable Diffusion Demo completely free | Subject to queueing/rate-limiting, no registration |
| Leonardo AI ↗ | 150 Fast Tokens/day, non-cumulative non-expiring | ≈25-37 base model images; third-party models (Kling/Sora etc.) burn credits at full speed |
| Adobe Firefly ↗ | 25 generative credits on first use, expire after 1 month | ≈25 standard images or 6 high-quality images |
| Ideogram ↗ | Free tier ~10 slow credits per day/week (significantly reduced from earlier this year) | Generated images default to public, cannot upload reference images |
| Google Gemini App (Nano Banana) ↗ | Free users ~20 images/day | Official: "limits change frequently, reset daily," no fixed number promised |
| Kuaishou Kling ↗ | Daily login grants 66 inspiration points, valid same day non-cumulative | ≈6 standard videos generatable |
| Alibaba Tongyi Wanxiang ↗ | After enabling Bailian, new users get ~50 text-to-image + 50s text-to-video, use within 90 days | Requires Alibaba Cloud identity verification |
Midjourney ↗ no longer has a free tier—Discord/web free trial was canceled in 2023; no credit card requirement but also no free quota. Specific numbers for Canva Magic Media and ByteDance Jimeng free quotas conflict across online sources; recommend logging in and testing yourself rather than copying numbers from any single tutorial.
Runway ↗: free tier 125 one-time credits (non-renewable non-recurring), export limited to 720p with watermark, does not include Gen-4 official model. Luma AI (Dream Machine) ↗: free tier ~30 generations/month, 5s each, 720p, no audio, with watermark. Pika: free tier limited to 480p with watermark; specific quota numbers vary by third-party source, recommend checking official site in real time.
OpenAI Sora: see §00—free quota canceled, product delisted; stop looking. Domestic Kling and Tongyi Wanxiang are listed in §06 (their video generation quotas share the same credit system as image generation).
| Product | Free Content | Notes |
|---|---|---|
| ElevenLabs ↗ | 10,000 credits/month (~10 min Multilingual v2 or 20 min Flash TTS) | Non-commercial use only, includes 3 Studio projects |
| Suno ↗ | 50 credits/day (non-cumulative, ~10 songs/day) | Limited to v4.5-all model, no commercial license; industry reports say free download permissions tightening from September |
| Udio ↗ | 10 credits/day + 100 credits/month (non-stackable non-cumulative) | Max 3 full-length songs/day, officially no credit card required |
| HeyGen ↗ | 3 videos/month, 720p with watermark | Officially no credit card required |
| Open-source Whisper ↗ | Free transcription via HuggingFace official Space | Subject to ZeroGPU daily quota limits (see §09) |
| iFlytek Online TTS ↗ | Personal developers 10K uses, enterprise developers 20K uses, 3-month validity | Includes 5 basic voice presets |
PlayHT has shut down (see §00)—acquired by Meta in July 2025, consumer product closed end of that year, official site inaccessible; skip any tutorial still recommending it. Tencent Zhiying's free quota couldn't be first-party verified; confirm by logging into your account.
| Platform | Free Quota | Notes |
|---|---|---|
| Google Colab ↗ | Single session max 12 hours, GPU/TPU model "varies over time" | No fixed weekly hours promised; community observes 15-30 hours/week fluctuation |
| Kaggle Notebooks ↗ | 30 hours GPU quota/week (official figure, still valid as of August 2026), Tesla P100 or T4×2 | Interactive session 60-min idle timeout |
| HuggingFace Spaces (ZeroGPU) ↗ | Free account 5 min/day, unauthenticated 2 min/day | Daily quota resets every 24 hours, max 2 ZeroGPU Spaces hosted |
| Lightning AI Studio ↗ | Free plan max 30 credits, 1 persistent Studio (restarts every 4 hours) | Free hours vary greatly by GPU type (T4 ~75 hours, H200 ~3 hours); recommend testing |
| Paperspace Gradient ↗ | Free subscription now only CPU (C4) and GPU (M4000) free instance types | Rumored free P5000 and other GPUs moved to paid plans only—this rumor is confirmed |
| Modal ↗ | $30 free compute credit/month, long-term valid | Credits don't carry over month-to-month; since August 2026, token-billed hosted model endpoints restricted to Team/Enterprise only |
No extra steps needed—register an account (some don't even require that) and start using; quota ceilings generally don't publish fixed numbers:
ChatGPT Free: official wording "daily text conversations unlimited," subject to abuse-protection limits; image generation/voice/file analysis billed separately. Claude.ai Free: rolling 5-hour window, specific message count not fixed/published, only Sonnet series available, Opus not open. Gemini App / Microsoft Copilot (Bing/Edge sidebar): free, hit ceiling and get "try later" prompt, no public hard number. Perplexity Free: standard web search basically unlimited, Pro Search has limited daily quota (3-5 per source variation). Meta AI: still free; paid subscription "Meta One" launched May 27, 2026 ↗, but official stance: "more casual users will continue to get it free."
GitHub Student Developer Pack ↗—after school email verification, bundles JetBrains full suite (free, ~$289/year value), DigitalOcean $200 credit, Azure $100 credit, plus newly added Camber (40 CPU-hours + 5 GPU-hours + 50 agent messages/month, connectable to Cursor/Claude Code via MCP). Copilot portion has changed dramatically: unlimited Copilot Pro for students canceled March 2026, replaced by 200 AI credits/month; new applications suspended April 20 (reason: excessive agent-mode consumption); existing users retain original benefits, new students currently get standard Copilot Free.
Azure for Students ↗: no credit card, $100 credit renewable once/year, applied against Azure OpenAI/Foundry pay-as-you-go (not an independent free quota).
Independent student verification channels are tightening across the board: Perplexity's ".edu one year free" ↗ expired January 2026, now Education Pro at half-price $10/month; Cursor student free Pro year closed to new users June 25, 2026 (see §00). Domestic Alibaba Cloud "Yun Gong Kai Wu" ↗ grants ¥300 voucher/year + 1-core 2GB ECS free 12 months after identity + Xuexin verification, but note that most "AI large model token quotas" from domestic cloud providers are universal benefits for all new users, not student-exclusive—don't conflate them with student programs.
Hackathon channels: there is no perennial AI credit pool anyone can self-serve at any time; the mechanism is event organizers apply first, then distribute credit codes to registered participants. OpenAI has an official application form ↗ for organizers; Anthropic's Claude Campus Program ↗ targets campus organizations to apply for credits to run events; Groq's own free developer tier is already generous (30,000 tokens/min), many hackathons just use the free tier; MLH partnership page ↗ offers $300 Gemini credits.
No institutional endorsement needed; individuals can self-serve apply:
AWS Activate Founders ↗: $1,000, no VC endorsement needed, just complete company documentation and ≤10 years since founding; credits directly convertible to third-party model usage on Bedrock. Modal: $30 free compute credit/month, long-term valid, included on registration. Google Cloud ↗: $300 trial for new users, 90-day validity, but cannot cover Gemini API's independent billing.
Requires institutional/academic endorsement; individual developers most likely can't get these—skip:
OpenAI Researcher Access Program ↗ ($1,000, requires academic institution endorsement, only 4 approval windows/year); Anthropic Claude for Startups ↗ (amount undisclosed, requires institutional investor funding + company ≤4 years); Together AI/Fireworks AI/Replicate large Startup Programs ($10K-$50K level, all require company identity).
Opera AI (formerly Aria): official page states "free to use without login or registration"; restructuring upgrade since October 2025 remains completely free; what actually costs money is the separate agentic browser Opera Neon at $19.90/month. Brave Leo: free tier requires no account, open and use directly; can call Llama 3.1, Qwen 3, Claude Haiku 3.5, GLM and other models; conversation history stored locally.
Beyond free, there's a category worth including: not $0, but cheap enough that "one free-ride lasts a long time" at standard pricing (excluding first-month promotional pricing). Conclusion first: the threshold (monthly fee ≤¥10 or ≤$3) is strict; image/video generation and cloud services currently have no qualifying standard-price options—everything found is $4+; not forcing entries. The solid finds are concentrated in coding IDEs and model API pay-as-you-go.
Coding IDEs—the only confirmed standard low-price plan: ByteDance Trae's paid entry tier Trae Lite, $3/month ↗, includes base quota + bonus usage, unlimited auto-completions, 2 concurrent cloud tasks; effective since February 2026 and unchanged. Compared to Cursor and Windsurf starting at $20/month, this is the only "long-term standard price" rather than "promotional price" low-cost plan.
Model API pay-as-you-go—converting to "how much does ¥10 buy" is more practical: per-token pricing varies significantly across providers; rather than memorizing free quotas, memorize unit-price conversions:
| Model | Input/Output price (per million tokens) | ≈What ¥10 RMB buys |
|---|---|---|
| Zhipu GLM-4.7-Flash | Free | No conversion needed, just use it |
| Alibaba Qwen-flash (≤128K context) | ¥0.15 / ¥1.5 | ~6.67M output tokens, cheapest of the four |
| DeepSeek-v4-flash (off-peak) | ¥1.5 / ¥4.5 (off-peak) | ~2.22M output tokens; peak price ¥3/¥9, half off requires timing |
| Kimi k2.6 | ¥6.5 (cache miss) / ¥27 | ~370K output tokens, most expensive of the four |
The conclusion is straightforward: the free GLM-4.7-Flash is itself the best value option; if you must pick a paid one, off-peak DeepSeek ↗ is several times more cost-effective than always-online Kimi ↗, and Qwen-flash ↗ is the cheapest paid option.
Borderline but not fully qualifying, for reference only: iFlytek Spark App and iFlytek Tingxie transcription continuous-pack standard renewal prices are ¥15-18/month (~$2-2.5), meeting the "$3" threshold but exceeding the "¥10" threshold, and first-month promos are often as low as ¥9.9 or ¥6—promotional prices don't count as long-term free-tier; always check the standard price before renewing; Alibaba Tongyi Qianwen's newly launched "Premium Member" in August 2026 is ¥19/month standard, also on the threshold edge—these two pieces of info ↗ couldn't be second-verified within the official app, suggest opening and confirming yourself before subscribing. Closest to the threshold in cloud services is DigitalOcean smallest instance at $4/month ↗, still no standard option truly below $3.
The previous two sections were about "how to spend less"; this section flips the calculation: if you're already using a coding subscription plan like Cursor or Claude Code, is it actually worth it, and is it cheaper than connecting to APIs yourself? The core variable is cache hit rate—each round of a coding agent request is typically "a large chunk of reused code context (should hit cache) + a small new instruction"; the real unit price must be weighted by hit rate, not just the base price listed on the website.
Subscription plans: who's quantifying publicly, who's deliberately obfuscating
| Product | Price | Usage disclosure method |
|---|---|---|
| GitHub Copilot | Pro $10 / Pro+ $39 / Max $100 | Only one dollarized: directly gives $15/$70/$200 credit quotas, deducts at each model's actual API price↗ |
| OpenAI Codex | Plus $20 / Pro $100(5x) / $200(20x) | Gives specific message count ranges (e.g. Plus 10-100 messages per 5-hour window on flagship model), estimates not hard caps↗ |
| Claude Code | Pro $20 / Max $100(5x) / $200(20x) | Only gives multipliers; officially states "no longer publishing token counts corresponding to multipliers"↗, dual limits of 5-hour window + weekly total; Claude Code shares usage pool with web version↗ |
| Cursor | Pro $20 / Pro+ $60 / Ultra $200 | Generous own model pool; "Other Models" pool deducted at official API prices, but dollar credit not labeled↗ |
| Gemini Antigravity | Pro $20 / Ultra $100(5x) / $200(20x) | Also only gives multipliers, not absolute numbers↗ |
| Windsurf | Free/$20/$200/Teams | Merged into Devin pricing page; even "how many credits consumed" disclosure abandoned, official wording: "varies by task"↗ |
No official source has published "this dollar amount equals how many API calls"—third-party estimates suggest Claude Code Max 20x heavy users' actual API-equivalent cost could reach $600-1,500/month, but this is community blog modeling, not official data (third-party estimate ↗).
API + Cache Pricing Panorama (per million tokens, USD)
| Model | Base Input | Output | Cache Hit Price | Hit Price / Base Price |
|---|---|---|---|---|
| Claude Opus 5 | $5.00 | $25.00 | $0.50 | 10% |
| Claude Sonnet 5 | $2.00 | $10.00 | $0.20 | 10% |
| GPT-5.6 Sol | $5.00 | $30.00 | $0.50 | 10% |
| GPT-5.6 Terra | $2.00 | $12.00 | $0.20 | 10% |
| Gemini 3 Pro | $2.00 | $12.00 | $0.20 | 10% |
| Gemini 3.8 Flash | $0.75 | $3.75 | $0.075 | 10% |
| Grok 4.6 | $2.00 | $6.00 | $0.50 | 25% |
| GLM-5.3 (overseas) | $1.40 | $4.40 | $0.26 | 19% |
| Kimi K2.6 | $0.95 | $4.00 | $0.16 | 17% |
| DeepSeek V4-Flash (off-peak) | $0.22 | $0.66 | $0.007 | 3% |
| DeepSeek V4-Pro (off-peak) | $0.66 | $1.98 | $0.022 | 3% |
Sources: Anthropic official pricing ↗ and cache docs ↗, OpenAI pricing page ↗, Google Gemini pricing page ↗, xAI pricing page ↗, Zhipu open platform ↗, Moonshot open platform ↗, DeepSeek pricing page ↗. DeepSeek's hit price is only ~3% of the miss price, the largest discount with the most transparent mechanism; Anthropic and OpenAI are the only two where "cache writes also incur a surcharge" (1.25× base), meaning you pay extra once before hitting.
How cache hit rate changes real cost: assuming 90% hit rate (typical for normal agentic coding sessions), effective input unit price = 0.9 × hit price + 0.1 × base price. By this formula, Claude Sonnet 5 / GPT-5.6 Terra / Gemini 3 Pro effective unit price all drop to $0.38 (19% of base), and DeepSeek V4-Flash (off-peak) drops to $0.028 (13% of base)—each additional 10 percentage points of hit rate benefits models like DeepSeek where "hit price = 3% of base price" the most.
Applied to one conversation round: assume a typical agent round is "50K tokens reused context (90% hit) + 1,500 tokens new output":
| Model | Cost per round | 300 rounds/month (light) | 1,500 rounds/month (heavy) |
|---|---|---|---|
| DeepSeek V4-Flash (off-peak) | $0.0024 | $0.72 | $3.6 |
| DeepSeek V4-Pro (off-peak) | $0.0073 | $2.2 | $10.9 |
| Gemini 3.8 Flash | $0.0128 | $3.8 | $19.2 |
| Kimi K2.6 | $0.018 | $5.4 | $26.9 |
| GLM-5.3 | $0.025 | $7.6 | $38.0 |
| Claude Sonnet 5 / GPT-5.6 Terra / Gemini 3 Pro | ~$0.034-0.037 | $10-11 | $51-56 |
| Claude Opus 5 | $0.085 | $25.5 | $127.5 |
| GPT-5.6 Sol | $0.093 | $27.8 | $138.8 |
Conclusion: Light users (~300 rounds/month) on Sonnet-class models pay only $10-11 in direct API costs—compared to Claude Code Pro at $20/month, pure pay-as-you-go is actually cheaper; the subscription buys "not worrying about bills + higher caps" peace of mind, not savings. Heavy users (1,500 rounds/month) see Sonnet-class API-equivalent cost rise to $51-56, at which point Claude Code Pro or Cursor Pro (both $20) are clearly better deals, and Max/Ultra tiers ($100-200) also start making sense. GitHub Copilot is the only one where you can precisely account: Pro's $15 credit buys ~440 rounds at Sonnet pricing or ~160 rounds at GPT-5.6 Sol—same dollar amount, choosing the cheaper model gets you nearly 3× more usage.
Heavy users on DeepSeek off-peak pricing pay only $3.6-10.9/month, an order of magnitude cheaper than any subscription plan—the tradeoff is connecting your own API and picking your own model, rather than lounging in a managed agent like Claude Code/Cursor. This loops right back to the logic from §03: self-built agents (Hermes Agent/Cline etc.) paired with cheap or free model APIs will always have the lowest cost ceiling—just requiring a bit more DIY effort.
The key to free-tier hunting is developing judgment, not treating some tutorial as a permanently valid instruction manual.
This article's closingPutting it all together, a basically zero-cost personal AI agent workflow looks like this: daily chat via ChatGPT/Claude.ai free web version; coding with Hermes Agent or Cline connected to OpenRouter's :free pool plus GLM-4.7-Flash as the main backend, adding $3 Trae Lite if budget allows; for web access, pair with Tavily or Firecrawl's 1,000 free calls/month; for code execution, use E) E2B's one-time $100 credit; for memory, use permanently free Mem0; deploy services on Cloudflare Workers—100K requests/day is enough for personal projects to run a long time; generate images with Leonardo AI (150 Fast Tokens/day) plus domestic' domestic Kling/Tongyi Wanxiang; for batch experiments, Kaggle Notebooks' 30 GPU hours/week plus Modal's $30/month basically suffice; for larger quotas, go through AWS Activate Founders' $1,000 self-serve application.
What you truly need to remember isn't any specific number, but the pattern the §00 table reminds you of: at least half the entries in this list will change within six months—free quotas get tightened (Brave Search API, Antigravity are examples), products get replaced (Gemini CLI, Bing Search API), former free deployment platforms become nominal-only (Railway, Fly.io). This pattern even played out once while writing this article: the time-limited stealth model Ox Alpha mentioned in the text was confirmed days later as Zhipu GLM-5.3-Flash in disguise, and the free route was delisted without warning. Coming back to verify periodically is more important than memorizing any single number.