# FetchAPI > A priced map of who sells data about websites and social platforms: thousands of API endpoints from > official APIs, resellers (RapidAPI, Apify, aggregators) and scraping services, each with its price > per call and per 1,000 items and the URL that states it. Ask it "what is the cheapest way to get > tracxn.com data?" and it answers with a ranked list of providers. This file is for LLMs and agents. Everything below is served from this site with no key and no login. ## Query it: GET /api/cheapest `GET /api/cheapest?site=` returns the cheapest sources of one website's data, one line per provider, cheapest first. JSON, CORS open, cached for an hour. GET /api/cheapest?site=tracxn.com GET /api/cheapest?site=pitchbook.com&limit=5 GET /api/cheapest?site=instagram.com&by=call&free=0 GET /api/cheapest # every site and platform covered, with its cheapest price `site` accepts a domain, a URL or a platform key (`x`, `instagram`, `tiktok`, `youtube`, `reddit`, `linkedin` ...). Parameters: | Parameter | Default | Meaning | |---|---|---| | `site` (or `domain`, `q`) | required | the website whose data you want | | `by` | `items` | sort by `items` ($ per 1,000 items returned) or `call` ($ per request) | | `limit` | 20 | 1-200 results | | `all` | off | `all=1` returns every endpoint, not just each provider's cheapest | | `capability` | none | narrow to one job, e.g. `profile`, `search_posts`, `other`, or a full key like `x.search_posts` | | `search` | none | match any words in endpoint name, capability, path or provider; useful for task phrases like `recent posts` | | `unpriced` | off | `unpriced=1` includes endpoints whose price is not published and $0 rows without a stated allowance | | `free` | on | `free=0` leaves out $0 rows (free tiers, samples) | Response fields, top level: - `site`, `label`: what the query resolved to. - `matched_by`, one of: - `site`: a site we have researched; - `platform`: a social platform or app store; - `name`: any other domain, matched by the site's name appearing in endpoint names, paths or vendor ids. - `search_applied`: true when `search=` produced usable task words and was applied to endpoint names, capabilities, paths and provider ids. - `providers`, `endpoints`: how many sell this site's data. - `official`: the site's own route, if we have it, with a link to the written analysis. - `results[]`: the ranked list (fields below). - `site_page`: a human-readable page with the same ranking. - `data_as_of`: the date of the newest price check. - `hint`: only present when nothing matched. It means the site has not been researched yet, not that nobody sells it. Fields of each `results[]` row: - `rank` - `provider`, `provider_name` - `official`: true when the row is the site's own route. - `free`: true when a zero-price endpoint has a stated free allowance. A $0 with no allowance is `zero-unverified`, excluded from default rankings, and included only with `unpriced=1`. - `status`: `paid`, `free`, `none`, or `zero-unverified`. - `free_allowance`: short published allowance label for verified free routes; inspect `free_tier` and verify its scope. - `verdict`: our read on the business (`large`, `real`, `early`, `unknown`, `pre-traction`, `cautionary`, `avoid`, `dead`). - `endpoint`, `capability`, `method`, `path` - `items_per_call` - `usd_per_1k_items {entry, volume}`, `usd_per_call {entry, volume}` - `confidence`: `stated`, `derived` or `estimated`. - `source`: the page that states the price. - `prices_checked`, `free_tier` - `provider_page` ## Generated site APIs: GET /api/site "Give it a URL, get an API." For a page we have generated an API for, `/api/site//?` fetches the live page and returns structured JSON. An LLM studied the page once and wrote a spec saying where the data lives: a JSON endpoint the page calls, JSON embedded in the HTML, or the HTML with CSS selectors. Calls run that spec directly, with no LLM. GET /api/site # every generated API, its endpoints and params GET /api/site/hn-algolia-com # one API GET /api/site/hn-algolia-com/search_stories?q=supabase # call it Calls use parse.bot's format. **A success** is `{"status": "success", "data": {...}, "meta": {...}}`: - **List endpoints:** `data` holds `count`, a named list (e.g. `raises`, `companies`) and each filter echoed. Paged endpoints add `page`, `has_more` and the site's total. - **`get_` endpoints:** `data` is the whole record. - **Values are raw:** money is in integer dollars, dates are `YYYY-MM-DD`, flags are booleans, and nested lists stay nested. - **`meta`** has the source URL, the fetch time, and whether the page needed the Bright Data unblocker. `GET /api/site/` returns the API's contract: for each endpoint, its params (type, example, and which endpoint's output supplies them), its return fields with descriptions, a verified sample, its checks and its health. **An error** is `{"status": "error", "error": {"type", "message"}}`: | Type | Status | Meaning | |---|---|---| | `invalid_input` | 422 | a bad parameter | | `stale_input` | 404 | an unknown id or slug, or no such endpoint | | `upstream_error` | 502 | the site returned an error | | `blocked` | 503 | a bot wall, or the API is marked blocked or unsupported | | `upstream_timeout` | 504 | the site took too long | | `extraction_failed` | 500 | the page changed and needs regenerating | The rate limit is 30 calls a minute per client; going over returns 429. Every call fetches the site live; responses are not cached. Before generating from a URL, check `/api/cheapest?site=&all=1` for existing paid routes and verified free allowances that may already supply the requested data. Add `unpriced=1` only when you also want rows with unknown prices or unverified $0 amounts. Verify a candidate's terms, current docs and account-specific availability; catalog prices do not prove the endpoint works with your credentials. New site-specific APIs are generated by a maintainer (`python3 scripts/gen_api.py ` in the repo), not through this endpoint. ## SEO data: /api/seo (DataForSEO, keyed) Every DataForSEO v3 API, through our domain. Get a key at `/keys` (self-serve, with a rate limit and a monthly $ cap) and send it as `x-api-key`. The same key works for the other providers it is enabled for, at `/api/gw//` (`GET /api/gw` lists them), and `GET /api/keys/me` shows its usage. GET /api/seo # catalogue: namespaces and priced endpoints GET /api/seo/appendix/user_data # free: balance and limits POST /api/seo/serp/google/organic/live/regular # body: [{"keyword": "...", "location_code": 2840, "language_code": "en"}] POST /api/seo/dataforseo_labs/google/ranked_keywords/live # body: [{"target": "example.com", "location_code": 2840}] - **Request bodies** are DataForSEO's task arrays, unchanged ([docs](https://docs.dataforseo.com/v3/)). - **Responses** are `{"status", "data": , "meta": {"cost_usd", "time_ms", ...}}`. Add `?raw=1` to get DataForSEO's JSON as is. - **Namespaces:** serp, keywords_data, dataforseo_labs, backlinks, domain_analytics, on_page, content_analysis, content_generation, merchant, app_data, business_data, ai_optimization, dataforseo_trends, appendix. - **Limits:** 60 calls a minute per key. Nothing is cached. ## MCP server and social data: /api/mcp, /api/gw/ Live X/Twitter, Instagram and Telegram data through one key, as an MCP server and as plain HTTP. Every endpoint was called for real and is documented with what one item is, the fields worth knowing and the answer we got: human docs at `/mcp`, the same for agents at `/mcp.md`, one provider per file at `/docs/.md`, machine-readable at `/gateway-catalog.json` (sample answers at `/gw-samples/.json`). claude mcp add --transport http fetchapi https://fetchapi.co/api/mcp --header "Authorization: Bearer gw_..." - **Tools:** `search_endpoints(query)` -> `get_endpoint(tool)` (params with working values, fields, real answer) -> `call_endpoint(tool, params)`; plus `list_providers()` and `get_usage()`. Searching and docs need no key; calls do. - **HTTP:** `GET /api/gw//?` with `x-api-key`, e.g. `/api/gw/rapidapi-twitter-api45/timeline.php?screenname=nasa`, `/api/gw/flashapi/ig/info_username?user=nasa`, `/api/gw/rapidapi-telegram-channel/channel/message?channel=telegram&limit=5`. - **Only working endpoints are listed:** each one returned data through this gateway in its latest live test (`node scripts/test_gateway.js`); one that stops working drops out of the catalogue and the MCP tools until it passes. - **Pagination:** an endpoint that pages has a `pagination` entry (in `get_endpoint` and `/gateway-catalog.json`): `next_from` is where the next cursor sits in the answer and `param` is where it goes on the next call. Over MCP, `call_endpoint(tool, params, pages=N)` follows it for you (N up to 5, each page one metered call) and returns `pages` plus `meta.next_cursor` / `meta.next_page` to continue later. Over HTTP, add the cursor yourself. - **X (Twitter) search, Top and Latest:** - `GET /api/gw/rapidapi-twitter-api45/search.php?query=openai&search_type=Latest` (or `Top`): the answer has `timeline` (20 posts) and `next_cursor`; the next page is `...&search_type=Latest&cursor=`. - `GET /api/gw/rapidapi-twitter-api47/v3/search?query=openai&type=Latest` (`type` is REQUIRED: `Top`, `Latest`, `People`, `Photos` or `Videos`, capitalised): the answer has `data` (20 items) and `pagination.nextCursor`; the next page is `...&type=Latest&cursor=`. - Keep the same sort on every page (a Top cursor does not page Latest), URL-encode the cursor, stop when it comes back empty or repeats. Tested 2026-10-06: 3 pages of each sort on both providers, 60 unique posts each. - **Brand -> creators:** `GET /api/brand-creators?brand=glossier&enrich=10` (or the MCP tool `find_brand_creators`) returns the creators who posted about or worked with a brand: Instagram posts that tag it (with Instagram's paid-partnership flag) and TikTok videos that mention it (with TikTok's branded-content label), per creator posts, paid partnerships, likes, comments, views, post links and, with `enrich`, followers and engagement rate. Several metered calls per lookup. - **Cost:** each call is metered at the provider's per-request price on its entry plan and counts toward the key's monthly cap. Write actions (post, follow, like) are listed in the docs but never exposed over MCP. ## How to read the prices - **entry** is the price on the smallest paid self-serve plan; **volume** is the largest plan with a public price. Sorting uses the lower of the two. - **$ / 1k items** normalises by page size (a call that returns 20 posts is worth more than one that returns 1); it is null when a provider does not publish page size. **$ / call** is what every provider states. - **confidence**: - `stated`: copied from the provider's page. - `derived`: arithmetic on stated numbers, e.g. plan price รท quota. - `estimated`: an assumption was needed, e.g. third-party quotes for a sales-only product. - **null means unknown**, never zero. "Price not published" is common for official enterprise data (PitchBook, Tracxn); those rows rank after every priced one. - Always open `source` before buying. Marketplace listings change price, break, or charge on a second event the headline does not show. The written analyses call out the traps we found, e.g. Apify actors whose dataset-item fee is nominal while the real charge sits on another event. ## Caveats an agent should pass on - **Licences.** Most resold copies of a site's data are scraped, and the site's terms usually forbid scraping and resale (Crunchbase, PitchBook, Tracxn, Similarweb, Semrush, BuiltWith all do). A cheap reseller is fine for internal research; republishing its output is the buyer's legal risk. - **Success rates.** Marketplace listings report them (RapidAPI's appear in the analyses). A $1/1k API that fails one call in three costs $1.50/1k. - **Coverage is research, not exhaustive.** Marketplaces always hold more listings than one search finds. The analyses say how each search was done. ## Other machine-readable data - `/data.js`: `window.DATA = {...}`. Directory of every vendor, the cheapest endpoint per vendor per capability, site summaries, and traffic/SEO metrics. - `/endpoints.js`: `window.ENDPOINTS = {cols, rows, vendors}`. Every endpoint as an array row (column names in `cols`), plus each vendor's pricing tiers. - Source repo: https://github.com/MovieTime99/agent-data. - `data/endpoints.csv`: every endpoint, flat. - `tools/*.md`: written analyses, e.g. `tools/tracxn.md`, `tools/pitchbook.md`, `tools/company-signals.md`. ## How the data is made 1. **Research.** Agents and people read provider pricing pages and write one YAML document per provider (format: `data/endpoints/_SCHEMA.md` in the repo). For a new site, `/find-cheapest ` runs three steps: - a mechanical sweep of Apify, RapidAPI and the TaskFuel catalogue (`scripts/find_cheapest.py`); - three research agents, covering the official route, the resellers, and scraping it yourself; - an ingest of everything they find. 2. **Database.** A Supabase Postgres database is the source of truth. Prices are append-only dated snapshots, so history accumulates; nothing computable (cheapest, rank, $/1k) is stored. 3. **Build.** `scripts/build.py` exports the database, computes rankings, and writes the files this site serves (`data.js`, `endpoints.js`). A site's endpoints are worked out from names (`scripts/sites.py`): - the site's name must appear as a word in the endpoint name, the path or the vendor id; - an endpoint named "not " is a recorded look-alike and is excluded; - the vendor whose id is the site's name is the official route. 4. **Serve.** Vercel serves the committed files; `/api/cheapest` is a function that reads the same files, so the API, the dashboard and the repo's scripts always agree. The API holds no credentials and cannot change data. To add a site, a maintainer runs `/find-cheapest ` in the repo (Claude Code), or `python3 scripts/ingest.py --taxonomy site ""` for one already covered by name.