web-search-and-content
25–48 of 492netintel-production-440c.up.railway.app
netintel-production-440c.up.railway.app
Fetch any public URL and extract structured metadata — Open Graph tags, Twitter Card tags, canonical URL, title, description, favicon, article author/date, JSON-LD structured data, and content type — so agents can preview, classify, and enrich links without building their own scraper.
netintel.dev
netintel.dev
Fetch and parse any XML sitemap or sitemap index file — returns URLs with their priority, change frequency, and last modified date, follows sitemap indexes one level deep (first 3 child sitemaps), and auto-discovers the sitemap from robots.txt or common paths when given a bare domain — so agents can enumerate site content for crawling, indexing, and SEO analysis.
netintel-production-440c.up.railway.app
netintel-production-440c.up.railway.app
Fetch and parse a domain's robots.txt file — returns all crawl rules by user-agent, sitemap URLs, crawl delay settings, and checks whether a specific path is allowed or blocked for any bot — so agents can respect crawl policies and locate sitemaps before scraping.
netintel-production-440c.up.railway.app
netintel-production-440c.up.railway.app
Fetch and parse any RSS 2.0 or Atom feed URL and return structured articles with title, link, description, author, publish date, and categories — so agents can monitor content sources, build news pipelines, and process any feed without building their own XML parser.
netintel.dev
netintel.dev
Fetch any public URL and extract structured metadata — Open Graph tags, Twitter Card tags, canonical URL, title, description, favicon, article author/date, JSON-LD structured data, and content type — so agents can preview, classify, and enrich links without building their own scraper.
screenshots.underscoredone.com
screenshots.underscoredone.com
Renders a webpage exactly as a real browser would, including all its scripts and dynamic content, then takes a picture of it. You can capture just the visible area or the whole scrollable page, choose the screen size, wait for slow-loading images to finish, and hide cookie banners or popups before the picture is taken.
gateway.spraay.app
gateway.spraay.app
Discover RTP robots. Filter by capability, chain, price, status.
api.webbersites.com
api.webbersites.com
Social share / OpenGraph checker: extracts og:*, twitter:*, title, description, canonical and robots meta from any URL, verifies the og:image actually loads and is a raster format, and returns problems, warnings, and a verdict. For publishing and SEO agents shipping pages that get shared.
websearch.use.x402atlas.com
websearch.use.x402atlas.com
Real-time web search for AI agents — search the web and get ranked SERP-style results as clean JSON: title, URL, snippet, relevance score, optional LLM-generated answer. Find up-to-date information online — latest news, current events, finance. Date, domain and exact-phrase filters.
websearch.use.x402atlas.com
websearch.use.x402atlas.com
Read any web page or article: fetch up to 5 URLs in one batch call and extract the main content as clean, LLM-ready markdown or plain text. Convert webpages to markdown, scrape page text, pull article content for RAG ingestion and AI agent context.
netintel.dev
netintel.dev
Detect the language of any text input using stopword-set matching and Unicode script analysis — returns the detected language, ISO code, a high/medium/low confidence level, and top alternative languages so agents can route multilingual content, trigger translation workflows, and classify text without a paid NLP API.
netintel.dev
netintel.dev
Text classification API — zero-shot text classifier / categorization: caller supplies 2–20 labels and Claude Haiku returns the best-matching category with confidence and per-label scores. Route, tag, triage, and detect intent/topic with your own taxonomy in one call.
netintel-production-440c.up.railway.app
netintel-production-440c.up.railway.app
Detect the language of any text input using character frequency analysis and n-gram pattern matching — returns the detected language, ISO code, confidence score, and top alternative languages so agents can route multilingual content, trigger translation workflows, and classify text without a paid NLP API.
netintel.dev
netintel.dev
Text summarizer / summarization API — condense text, Markdown, or a URL into a TL;DR plus key bullet points using Claude Haiku. Accepts raw text or a URL (fetched and extracted), enforces an input cap, and returns a clean structured summary so agents can compress documents and articles in one call.
web-scraper-api-production-bf20.up.railway.app
web-scraper-api-production-bf20.up.railway.app
Fetch a URL and extract clean main-content text with title, description, word count, and char count; boilerplate (nav/footer/sidebar/ads) is stripped.
web-scraper-api-production-bf20.up.railway.app
web-scraper-api-production-bf20.up.railway.app
Extract structured elements: heading hierarchy (h1-h6), lists, tables as 2D arrays, and images with alt text, plus element counts.
api.craigmbrown.com
api.craigmbrown.com
Fast real-time news scan across 44+ curated domains (configs/search_domain_profiles.json), every item source-attributed. Settlement proof: ProofOfSettledOutcome (kind 30120, data/proof_settled_outcomes.jsonl). The quick-scan tier below research.topic-deep-researcher's slower, citation-grade report.
x402.ottoai.services
x402.ottoai.services
Neural web search for agents (Exa upstream): pass a query, get the most relevant live web pages with title, URL, published date, author, relevance score and a clean text snippet as structured JSON. The find-then-read pair for /web-extract — RAG retrieval, research, fact-finding, monitoring. Keyless, pay-per-call, cached by query. A no-match query returns an empty result set as a valid answer.
2s.io
2s.io
Render a URL in a real headless browser (JavaScript executed) and return its article content with the clutter stripped — same cleaning + output formats as /api/url/clean, but for client-rendered / SPA pages whose content only appears after JS runs (where a raw HTTP fetch returns an empty shell). `format`: markdown (default), text, both (JSON envelope), html (self-contained reader page, raw text/html), or pdf (typeset reading doc, raw application/pdf). Optional `waitUntil` (load|domcontentloaded|networkidle0|networkidle2, default networkidle2) and `timeoutMs` (1000-15000, default 12000) control how long to let JS settle. Use /api/url/clean instead for server-rendered pages — it is faster and cheaper; only reach for render when JS rendering is required.
2s.io
2s.io
AI web search optimized for agents. Returns ranked results with the relevant extracted content of each page (not just a link + blurb), plus a relevance score. topic=news for recent reporting. Distinct from search.web (raw SERP) — this returns clean, LLM-ready page content per result.
text-classifier.api.klymax402.com
text-classifier.api.klymax402.com
Classify text content into categories with confidence scores and readability metrics
visual.hugen.tokyo
visual.hugen.tokyo
Extract structured text, tables, and metadata from any PDF URL. Returns page-by-page text content, detected tables as JSON arrays, and document metadata (title, author, page count). No PDF library or OCR setup needed — pay per extraction. AI agent API for document intelligence and data extraction
netintel.dev
netintel.dev
Fetch and parse a domain's robots.txt file — returns all crawl rules by user-agent, sitemap URLs, crawl delay settings, and checks whether a specific path is allowed or blocked for any bot — so agents can respect crawl policies and locate sitemaps before scraping.
scout.hugen.tokyo
scout.hugen.tokyo
Search ArXiv for academic papers and preprints in AI, machine learning, CS, and mathematics. AI agent API for scientific research. Accepts USDC payments on Base and Solana