Scrape any web page into structured JSON data with ScrapeNinja and AI
n8n-workflow-2812.n8n.io
· n8n.io
Disclaimer: This template only works on self-hosted for now, as it uses a community node. Use Case Web scrapers often break due to web page layout changes. This workflow attempts to mitigate this problem by auto-generating web scraping data extractor code via LLM. How It Works This workflow leverages ScrapeNinja n8n community node to: - scrape webpage HTML, - feed it into LLM (Google Gemini) and ask to write a JS extractor function code, then it - executes the written JS extractor against scraped HTML to extract useful data from webpage (the code is safely executed in a sandbox) Installation To install ScrapeNinja n8n node, in your self-hosted instance, go to Settings -> Community nodes, enter "n8n-nodes-scrapeninja", and install. Make sure you are using at least v0.3.0. See this in action: https://www.linkedin.com/feed/update/urn:li:activity:7289659870935490560/
n8n-workflow-2812.n8n.io via a single DNS TXT record to add the
verified by owner badge, embed an Agenstry badge on your README, and earn back the missing conformance points listed below.
Dispute or improve this rating
F
Conformance score: 19/100
F-grade: card is reachable but fails most operational signals.
click to expand breakdown ▾
click to collapse breakdown ▴
agent-card.json changed within the last 7 days. We track these so downstream callers can react.
Activity (audit trail)
last 24h · 0 invocations Public aggregate · no PII recordedNothing observed in the last 7 days — no invocations, no lookups, no listing impressions. Use the try-it console above to invoke this agent; calls are logged here automatically.
Endpoints
| Agent card | https://n8n.io/workflows/2812 |
| Provider | https://n8n.io |
| Docs | https://n8n.io/workflows/2812 |
Health · last 0 probes
Similar agents embedding-nearest
Embed your Agenstry badge
Paste any of these into your README, agent card, or marketing page. Each badge auto-updates and links back to this page.
Markdown / HTML snippets
[](https://agenstry.com/agents/n8n-workflow-2812.n8n.io) [](https://agenstry.com/agents/n8n-workflow-2812.n8n.io) [](https://agenstry.com/agents/n8n-workflow-2812.n8n.io) [](https://agenstry.com/agents/n8n-workflow-2812.n8n.io)
Audit-grade evidence bundle
JSON snapshot for vendor-review files. Add ?sign=true for a JWS-signed envelope verifiable against
our JWKS. See the methodology.
Raw agent card JSON
{
"_source": "n8n.io",
"workflow": {
"id": 2812,
"name": "Scrape any web page into structured JSON data with ScrapeNinja and AI",
"totalViews": 80527,
"price": null,
"purchaseUrl": null,
"recentViews": 2,
"createdAt": "2025-01-28T07:07:34.371Z",
"user": {
"username": "scrapeninja",
"verified": true
},
"readyToDemo": null,
"nodes": [
{
"id": 838,
"icon": "fa:mouse-pointer",
"name": "n8n-nodes-base.manualTrigger",
"codex": {
"data": {
"resources": {
"generic": [],
"primaryDocumentation": [
{
"url": "https://docs.n8n.io/integrations/builtin/core-nodes/n8n-nodes-base.manualworkflowtrigger/"
}
]
},
"categories": [
"Core Nodes"
],
"nodeVersion": "1.0",
"codexVersion": "1.0"
}
},
"group": "[\"trigger\"]",
"defaults": {
"name": "When clicking \u2018Execute workflow\u2019",
"color": "#909298"
},
"iconData": {
"icon": "mouse-pointer",
"type": "icon"
},
"displayName": "Manual Trigger",
"typeVersion": 1,
"nodeCategories": [
{
"id": 9,
"name": "Core Nodes"
}
]
},
{
"id": 1123,
"icon": "fa:link",
"name": "@n8n/n8n-nodes-langchain.chainLlm",
"codex": {
"data": {
"alias": [
"LangChain"
],
"resources": {
"primaryDocumentation": [
{
"url": "https://docs.n8n.io/integrations/builtin/cluster-nodes/root-nodes/n8n-nodes-langchain.chainllm/"
}
]
},
"categories": [
"AI",
"Langchain"
],
"subcategories": {
"AI": [
"Chains",
"Root Nodes"
]
}
}
},
"group": "[\"transform\"]",
"defaults": {
"name": "Basic LLM Chain",
"color": "#909298"
},
"iconData": {
"icon": "link",
"type": "icon"
},
"displayName": "Basic LLM Chain",
"typeVersion": 2,
"nodeCategories": [
{
"id": 25,
"name": "AI"
},
{
"id": 26,
"name": "Langchain"
}
]
},
{
"id": 1262,
"icon": "file:google.svg",
"name": "@n8n/n8n-nodes-langchain.lmChatGoogleGemini",
"codex": {
"data": {
"resources": {
"primaryDocumentation": [
{
"url": "https://docs.n8n.io/integrations/builtin/cluster-nodes/sub-nodes/n8n-nodes-langchain.lmchatgooglegemini/"
}
]
},
"categories": [
"AI",
"Langchain"
],
"subcategories": {
"AI": [
"Language Models",
"Root Nodes"
],
"Language Models": [
"Chat Models (Recommended)"
]
}
}
},
"group": "[\"transform\"]",
"defaults": {
"name": "Google Gemini Chat Model"
},
"iconData": {
"type": "file",
"fileBuffer": "data:image/svg+xml;base64,PHN2ZyB4bWxucz0iaHR0cDovL3d3dy53My5vcmcvMjAwMC9zdmciIHhtbG5zOnhsaW5rPSJodHRwOi8vd3d3LnczLm9yZy8xOTk5L3hsaW5rIiB2aWV3Qm94PSIwIDAgNDggNDgiPjxkZWZzPjxwYXRoIGlkPSJhIiBkPSJNNDQuNSAyMEgyNHY4LjVoMTEuOEMzNC43IDMzLjkgMzAuMSAzNyAyNCAzN2MtNy4yIDAtMTMtNS44LTEzLTEzczUuOC0xMyAxMy0xM2MzLjEgMCA1LjkgMS4xIDguMSAyLjlsNi40LTYuNEMzNC42IDQuMSAyOS42IDIgMjQgMiAxMS44IDIgMiAxMS44IDIgMjRzOS44IDIyIDIyIDIyYzExIDAgMjEtOCAyMS0yMiAwLTEuMy0uMi0yLjctLjUtNCIvPjwvZGVmcz48Y2xpcFBhdGggaWQ9ImIiPjx1c2UgeGxpbms6aHJlZj0iI2EiIG92ZXJmbG93PSJ2aXNpYmxlIi8+PC9jbGlwUGF0aD48cGF0aCBmaWxsPSIjRkJCQzA1IiBkPSJNMCAzN1YxMWwxNyAxM3oiIGNsaXAtcGF0aD0idXJsKCNiKSIvPjxwYXRoIGZpbGw9IiNFQTQzMzUiIGQ9Im0wIDExIDE3IDEzIDctNi4xTDQ4IDE0VjBIMHoiIGNsaXAtcGF0aD0idXJsKCNiKSIvPjxwYXRoIGZpbGw9IiMzNEE4NTMiIGQ9Im0wIDM3IDMwLTIzIDcuOSAxTDQ4IDB2NDhIMHoiIGNsaXAtcGF0aD0idXJsKCNiKSIvPjxwYXRoIGZpbGw9IiM0Mjg1RjQiIGQ9Ik00OCA0OCAxNyAyNGwtNC0zIDM1LTEweiIgY2xpcC1wYXRoPSJ1cmwoI2IpIi8+PC9zdmc+"
},
"displayName": "Google Gemini Chat Model",
"typeVersion": 1,
"nodeCategories": [
{
"id": 25,
"name": "AI"
},
{
"id": 26,
"name": "Langchain"
}
]
}
]
},
"detail": {
"description": "Disclaimer: This template only works on self-hosted for now, as it uses a community node. Use Case Web scrapers often break due to web page layout changes. This workflow attempts to mitigate this problem by auto-generating web scraping data extractor code via LLM. How It Works This workflow leverages ScrapeNinja n8n community node to: - scrape webpage HTML, - feed it into LLM (Google Gemini) and ask to write a JS extractor function code, then it - executes the written JS extractor against scraped HTML to extract useful data from webpage (the code is safely executed in a sandbox) Installation To install ScrapeNinja n8n node, in your self-hosted instance, go to Settings -> Community nodes, enter \"n8n-nodes-scrapeninja\", and install. Make sure you are using at least v0.3.0. See this in action: https://www.linkedin.com/feed/update/urn:li:activity:7289659870935490560/"
}
}