Docs

domsect.io API — Raw Compute. Pure Signal.

The high-throughput DOM pruning and structured web extraction API. Drop-in JSON from any URL. Base URL: https://api.domsect.io

Authentication

Bearer token in the Authorization header. Sign in with GitHub for 10 free credits, or apply for the 500-credit sandbox on the homepage — no card.

POST /v1/extract

Send a URL plus a field-type schema. Routing, HTML denoising and model selection all happen server-side — there is no model parameter.

urlstring, required — the public page to extract.
schemaobject, required — field name → type (string, number, boolean, array, object).
stealthboolean, default false — Tier-2 stealth extraction for heavily protected sites. Coming soon; 3 credits/page when live.
us_soilboolean, default false — US-Soil Dedicated routing. Guaranteed zero cross-border data transfer: DOM pruning, intermediate payloads, and LLM structured extraction are executed exclusively on US-domestic compute clusters. Offshore failover is strictly disabled — if domestic capacity is temporarily unavailable, requests fail-fast at 0 credits rather than rerouting overseas. Billed at 2× credits.
curl -X POST https://api.domsect.io/v1/extract \
  -H "Authorization: Bearer *********" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://example.com/product/123",
    "schema": { "title": "string", "price": "number", "in_stock": "boolean" },
    "stealth": false,
    "us_soil": false
  }'

Response

{
  "data": { "title": "...", "price": 49.99, "in_stock": true },
  "latency_ms": 2140,
  "usage": { "credits": 1 }
}

data is shaped exactly like your schema. usage.credits tells you what this call cost — 1 for Tier-1, 3 for Tier-2 (when live) — doubled when us_soil is true.

Data-sovereignty boundary: the zero-cross-border guarantee covers DomSect's internal compute and processing infrastructure. Fetching the publicly accessible target URL you submit travels over the public internet and is subject to that target domain's hosting location and routing.

Billing

Credits deduct only on HTTP 200 with schema-valid JSON. 404s, blocks, timeouts and failed fetches cost 0 credits. See /pricing for plans — about half the per-page cost of Firecrawl.

Errors

400Bad request — missing url or malformed schema.
401Missing or invalid API key.
404Page not found or unfetchable. Costs 0 credits.
422Extraction failed to produce schema-valid output. Costs 0 credits.
429Rate limit hit — back off and retry.
5xxGateway or model failure. Costs 0 credits.

Custom Crawler Tailoring

Custom Crawler Engineering for Complex Targets

While our core /v1/extract pipeline automatically strips DOM noise and guarantees strict JSON schemas across the public web, heavily protected platforms (Cloudflare Turnstile, AWS WAF, DataDome) may block generic headless fetchers.

If your production pipeline requires custom scraper tuning for specific high-friction sites, our verified engineer network can assist with bespoke automation scripts that route directly into your DomSect API endpoint.

Service Policy & Inquiries

Eligibility — Strictly reserved for active paying subscribers (Starter, Growth, Business). Requests from Free Trial / Sandbox keys are automatically declined.
Scope — Limited strictly to Lawfully Publicly Available Information (PAI). We do NOT build tools to bypass authenticated logins, paywalls, or sensitive personal data (PII).
How to Request — Email [email protected] with the exact subject line:
[Custom Scraper Request] - Active Subscriber: <Your Account Email>
Please include sample target URLs and your required JSON schema in the email. Our engineering partners will evaluate technical feasibility and provide a direct delivery estimate.

Compliance boundary

URL-only, public data only — no logins, no paywalls, no private documents. Sensitive hits are blocked at the gateway at zero charge. Full terms: ToS · AUP.