Skip to main content

Firecrawl Tutorial 2026: Turn Any Site Into LLM-Ready Markdown

Firecrawl tutorial for 2026: get an API key, scrape a page to markdown, map and crawl a site, extract JSON in Python, and budget credits and rate limits.

Founder of IImagined.ai

Published
Oct 11, 2026
Reading time
10 min read
Quick answer

Firecrawl turns a URL into markdown or structured JSON with one API call. This tutorial runs one product-catalogue job through scrape, map, batch scrape, crawl and JSON mode in Python, then works out the credits: a plain page costs 1 credit, JSON mode costs 5, and the free plan includes 1,000 credits a month.

This Firecrawl tutorial takes one job, turning a store's product pages into data an LLM can use, through the five calls you need: scrape one page to markdown, map the site's URLs, batch scrape or crawl the pages, and extract fields with JSON mode. On the v2 API a plain page costs 1 credit and a JSON extraction costs 5, the free plan includes 1,000 credits a month, and the Python SDK installs with pip install firecrawl-py.

Checked October 2026 against Firecrawl's docs for scrape, map, crawl, JSON mode, Agent, billing and rate limits, plus the pricing page. This page is built from that documentation, not from benchmark runs: the code follows the official examples, and every credit total is arithmetic on the published rates.

By the end you will have a short Python script that lists a site's product URLs, pulls each page as clean markdown, and returns name, price and stock as typed JSON, plus a credit budget you can check before you run it. If you are still choosing between an API like this and writing your own Playwright scraper, start with our web scraping automation guide, which covers the build-it-yourself route.

The product-data job, end to end
  1. 01
    Scrape one page

    Read the markdown and confirm the data you want is in it.

  2. 02
    Map the site

    One call returns the URL list. Filter it to product pages.

  3. 03
    Batch scrape

    Send the filtered list as one job and poll for documents.

  4. 04
    Extract JSON

    Add a schema so each page returns typed fields.

  5. 05
    Store and use

    Save markdown for retrieval and JSON for your database.

What you need before you start

Nothing here needs a paid plan. The free plan's 1,000 credits cover every step in this tutorial several times over, as long as you keep the page limits small while you test.

Pre-flight check
  • Python 3 with pip, or Node if you prefer the JavaScript SDK
  • A Firecrawl account and an API key that starts with fc-
  • One product URL from the target site to test with
  • The site's terms and robots.txt read, so you know what you may collect
  • A list of the fields you want back: name, price, currency, stock
  • A page limit for the first run (10 is plenty)

Step 1: get a Firecrawl API key and install the SDK

Sign up at firecrawl.dev, open the dashboard and copy your key. Then install the Python package and put the key in an environment variable so it never lands in your code:

pip install firecrawl-py
export FIRECRAWL_API_KEY="fc-YOUR-API-KEY"

Expected result: pip show firecrawl-py prints a version. The SDK reads FIRECRAWL_API_KEY from the environment, or you can pass api_key when you create the client. Node users install with npm install firecrawl, and there is a CLI (npm install -g firecrawl, then firecrawl login) if you want to test from a terminal first.

Step 2: a Firecrawl Python example that scrapes one page

Start with a single page. One scrape tells you whether the content you care about survives the trip to markdown, and it costs 1 credit.

from firecrawl import Firecrawl

firecrawl = Firecrawl(api_key="fc-YOUR-API-KEY")

doc = firecrawl.scrape(
    "https://example.com/products/trail-shoe",
    formats=["markdown"],
)
print(doc.markdown)

Expected result: the page body as markdown, with navigation and footer stripped. Read it. If the price or the stock status is missing, the fix is here, not later: the content may sit outside the main content area, or load after a delay.

Scrape has sensible defaults, and most of the tuning is knowing which ones to change. These five come from the API reference:

OptionDefaultWhy you would change it
formats["markdown"]Ask only for what you will use; JSON, question and highlights add credits
onlyMainContenttrueStrips navigation, headers and footers; turn off if the data lives there
maxAge172800000 ms (2 days)A cached copy younger than this is returned; set 0 for a fresh fetch
proxyautoRetries with enhanced proxies if basic fails, at no extra credit
timeout60000 msQueue time counts against it, so raise it for big batches

One thing the cache does not do is save credits. Firecrawl's docs are explicit that a cached result still costs 1 credit per page; caching makes the response faster, nothing more.

Step 3: map the site and pick the pages

Map returns the URLs on a site without fetching their content. It draws mainly on the sitemap, and it costs 1 credit per call however many URLs come back, which makes it the cheapest way to find out what you are dealing with.

res = firecrawl.map(url="https://example.com", limit=500, sitemap="include")

product_urls = [link.url for link in res.links if "/products/" in link.url]
print(len(product_urls))

Expected result: a count of product URLs. The docs warn that map trades completeness for speed, so if the count looks low compared with the store's own category pages, use crawl for discovery instead. For a quick filter on the server side, map also accepts a search parameter that orders results by relevance to a term.

Step 4: batch scrape the list, or crawl the section

You now have two ways to fetch many pages. Batch scrape takes the list you just built. Crawl takes a starting URL and finds the pages itself. Batch is the predictable one, because you decide exactly which URLs get fetched.

job = firecrawl.batch_scrape(
    product_urls[:10],
    formats=["markdown"],
    poll_interval=2,
    wait_timeout=120,
)
for doc in job.data:
    print(doc.metadata.status_code, doc.metadata.source_url)

Expected result: ten lines, each with a status code and a URL. Check the status codes. Firecrawl reports two layers: the API call can succeed while the target page returned a 403 or 404, and that page still comes back as a document and still costs 1 credit. Stop retrying URLs that keep failing.

Crawl suits the case where you want everything under a path and do not trust the sitemap. The path filters are regular expressions on the URL pathname:

curl -s -X POST "https://api.firecrawl.dev/v2/crawl" \
  -H "Authorization: Bearer $FIRECRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://example.com",
    "limit": 100,
    "includePaths": ["^/products/.*"],
    "scrapeOptions": { "formats": ["markdown"], "onlyMainContent": true }
  }'

The response is a job ID. Poll GET /v2/crawl/<id> until the status is completed, or use the SDK's crawl method, which waits and handles pagination for you. Results stay available through the API for 24 hours. If you would rather be told than poll, crawl can post crawl.page and crawl.completed events to a webhook; our n8n webhook trigger guide shows how to catch one.

Step 5: extract structured product data with JSON mode

Markdown is right for retrieval and summarising. For a database row you want fields. JSON mode runs an LLM over the page and returns data that matches a schema you define, in the same scrape call.

from pydantic import BaseModel

class Product(BaseModel):
    name: str
    price: float
    currency: str
    in_stock: bool

result = firecrawl.scrape(
    "https://example.com/products/trail-shoe",
    formats=[{
        "type": "json",
        "schema": Product.model_json_schema(),
        "prompt": "Extract the product shown on this page.",
    }],
)
print(result.json)

Expected result: a dictionary with the four fields. In the v2 API the schema sits inside the format object; the older jsonOptions parameter from v1 no longer exists, which is the usual reason a copied snippet fails. You can also pass only a prompt and let the model choose the structure, but a schema is what keeps 500 pages consistent.

Two alternatives are worth knowing. For product pages specifically, the product format returns title, price, availability and variants without an LLM call or a schema, and the billing table lists no surcharge for it. And when you do not know which URLs hold the answer, the Agent endpoint takes a prompt and finds the pages itself; Firecrawl describes it as the successor to the older /extract endpoint, which still works.

Which Firecrawl call fits the job
URLs known
Scrape. Add JSON mode or the product format when you want fields instead of markdown.
Batch scrape the list. Same options as scrape, one job to poll, cost known in advance.
URLs unknown
Search, or Agent with a prompt, when you do not know where the answer lives.
Map to list URLs for 1 credit, or crawl to discover and scrape in one job.
One page
Many pages
ModeYou give itYou get backCredits
ScrapeOne URLMarkdown, HTML, links, screenshot or JSON for that page1 credit per page
Batch scrapeA list of URLs you already haveOne job to poll, a document per URL1 credit per page
MapA site rootA list of URLs with titles, no page content1 credit per call
CrawlA start URL plus a limitEvery page it discovers and scrapes1 credit per page
JSON modeA scrape or crawl with a schema or promptFields extracted by an LLM+4 credits per page
AgentA prompt, URLs optionalData it searches for and gathers itself5 free runs a day, then usage-based

Firecrawl API credit costs for the whole job

Credits are the same on every plan, so the budget is multiplication. Take an illustrative catalogue of 500 product pages. Mapping it costs 1 credit. Scraping all 500 to markdown costs 500. Running JSON mode on all 500 costs 2,500, because each page is 1 base credit plus 4 for the extraction.

Credits for an illustrative 500-page catalogue, by mode
Map the site (1 call)
1
Scrape to markdown
500
Product format
500
JSON mode
2,500

Source: Firecrawl billing docs, checked October 2026. The 500 pages are an illustrative input, not a measured run.

That arithmetic decides the plan. The free plan's 1,000 credits cover 200 JSON extractions or 1,000 plain pages a month. Hobby's 5,000 cover 1,000 JSON pages. If your fields are standard product fields, the product format has no listed surcharge, which would put the same 500 pages at a fifth of the JSON-mode cost; confirm it on your first ten pages by reading the job's credits used. And once markdown is in your own storage, the remaining spend is the model you send it to; our LLM API pricing comparison has the per-token rates.

Firecrawl rate limits and the settings that matter

Two separate limits apply. Requests per minute are counted per endpoint and per team, and going over returns a 429. Concurrent browsers cap how many pages render at once; extra jobs wait in a queue instead of failing.

PlanPrice (checked October 2026)Credits a monthConcurrent browsersScrape requests a minuteCrawl requests a minute
Free$01,0002102
Hobby$16 billed yearly, $19 monthly5,000510020
Standard$83 billed yearly100,00025500100
Growth$333 billed yearly500,000505,0001,000
Scale$599 billed yearly1,000,00010010,0002,000

On the free plan, 10 scrape requests a minute means 500 single scrape calls take at least 50 minutes. That is the practical argument for batch scrape and crawl: one request starts the whole job, and Firecrawl works through it at your concurrency. The older extract endpoint shares Agent's limit.

A safe run order for a first job
  1. 1
    Scrape one page

    Read the markdown. Fix content problems before you spend more.

  2. 2
    Map and filter

    One credit. Check the URL count against what the site shows.

  3. 3
    Batch ten pages

    Look at every status code, not just the first document.

  4. 4
    Add the schema

    Run JSON mode on the same ten and compare with the pages by eye.

  5. 5
    Set the limits

    Pass limit on crawls, and maxConcurrency if the target site is small.

  6. 6
    Run the full list

    Save results within 24 hours, before the job expires.

Be a polite guest. Crawl's delay option waits between pages and forces concurrency to 1, and maxConcurrency caps parallel scrapes below your plan's ceiling. A small store's server does not need 25 browsers at once.

Firecrawl Docker: can you self-host it instead?

Yes. Firecrawl is open source under AGPL-3.0, and the repository includes a Docker Compose setup:

git clone https://github.com/firecrawl/firecrawl.git
cd firecrawl
# create .env first (see the self-host guide), then:
docker compose up --build -d

The API then answers on http://localhost:3002, and the same /v2/scrape calls work against it. The self-host guide is clear about the gaps: the default stack has no Fire-engine, so no advanced anti-bot handling, no screenshots and no page actions; Agent is a cloud feature; and JSON mode needs an LLM provider you wire up yourself. The quickstart also runs with API authentication off, so keep it off the public internet.

Troubleshooting: when a Firecrawl call fails

  • HTTP 402. You are out of credits, or a crawl's limit is more than your balance allows. Lower the limit or wait for the monthly reset.
  • HTTP 429. You passed the requests-per-minute limit for that endpoint. Back off, or move single scrapes into one batch job.
  • Markdown is missing the data. Set only_main_content=False, or add a wait if the content loads late.
  • Stale prices. The default cache window is two days. Pass max_age=0 for a fresh fetch.
  • JSON fields come back empty. Check that the markdown contains them first; the extraction can only read what the scrape captured.
  • A v1 snippet throws an error. Move the schema into formats as a json object and use the Firecrawl class.

Scraped data is only useful once something reads it. If you want an agent to call Firecrawl directly, it ships an MCP server, and our Model Context Protocol guide explains how that connection works. If you would rather call a model yourself, the Claude API tutorial picks up where this script ends, and the Supabase tutorial covers where to keep the rows.

Turning that pipeline into something customers log in to is a longer job. Our AI SaaS Builder program covers the part this tutorial stops at: the app, the database, auth and billing around the data.

Firecrawl tutorial: FAQ

What is Firecrawl used for?

Firecrawl is a web data API that turns pages into formats a language model can read. You send a URL and get clean markdown, HTML, links, a screenshot or structured JSON back, with JavaScript rendering, proxies and caching handled for you. Builders use it to feed retrieval systems, monitor competitor pages, enrich leads and give AI agents a way to read the live web.

How do I get a Firecrawl API key?

Create a free account at firecrawl.dev and copy the key from the dashboard; keys start with fc-. Pass it to the SDK as api_key or set the FIRECRAWL_API_KEY environment variable. Firecrawl's docs also describe a keyless tier for scrape, search and interact that is capped per IP address per day, which is enough to test one call before you sign up.

Is Firecrawl free?

There is a free plan with 1,000 credits a month, two concurrent browsers and no card required, checked October 2026. Credits do not roll over and there is no pay-as-you-go on the free plan, so requests return HTTP 402 once the month's credits are used. The open-source version can also be self-hosted with Docker at no licence cost, with fewer features.

How many credits does a Firecrawl scrape cost?

One credit per page for scrape and for each page a crawl reaches, according to Firecrawl's billing docs checked October 2026. Map costs one credit per call however many URLs it returns. JSON mode adds four credits per page, so a structured extraction costs five. PDF parsing adds one credit per PDF page, and search costs two credits per ten results.

What is the difference between scrape, crawl and map in Firecrawl?

Scrape fetches one URL and returns its content. Map returns a list of a site's URLs without fetching their content, mostly from the sitemap, so it is fast and costs one credit. Crawl starts at a URL, follows links and scrapes every page it reaches up to your limit. Use map plus batch scrape when you want to choose pages, and crawl when you want everything under a path.

Can I run Firecrawl with Docker?

Yes. The repository ships a Docker Compose setup: clone it, create a .env file with database settings and USE_DB_AUTHENTICATION=false, and run docker compose up --build -d. The API then listens on localhost:3002. Firecrawl's self-host guide says the default stack lacks Fire-engine, screenshots, page actions and the Agent endpoint, and that LLM extraction needs a model provider you connect yourself.

Does Firecrawl work with Python and Node?

Both have official SDKs. Python installs with pip install firecrawl-py and imports as from firecrawl import Firecrawl; Node installs with npm install firecrawl. The method names match the endpoints: scrape, crawl, map, search, batch scrape and agent. There is also a firecrawl command-line tool and an MCP server, so coding agents such as Claude Code can call it as a tool.

All Access · all four programs · $99/mo

Clean data is the easy half. The product around it is the work.

AI SaaS Builder, included in All Access, takes a pipeline like this one through the Claude API, Supabase, Next.js and payments until people can sign up for it, with the other three programs, live coaching and the private community in one subscription.

Start All Access — $99/mo →30-day money-back guarantee
Free · no signup

Building a scraper this week?

Join the free Discord to compare schemas, crawl settings and credit budgets with other builders.