Web Extraction Reference
Use useWebExtract to extract clean markdown text from public web pages for display, analysis, summarization, and other AI processing workflows.
useWebExtract
Returns
extract(function): Extract content from a URL.extractedText(string | null): Extracted markdown text.setExtractedText((text: string | null) => void): Update or clear the text locally.isLoading(boolean): Whether extraction is in progress.error(Error | null): Current extraction error.clearError(() => void): Clear the current error.
extract Input
url(string, required): URL of the public web page to extract.
extract resolves to markdown text on success or undefined on failure.
Basic Example
Import the hook from the web extraction module:
import { Markdown } from '@/components/ui/markdown'
import { useWebExtract } from '@/hooks/use-web-extract'
export default function App() {
const { extract, extractedText, isLoading, error } = useWebExtract()
const [url, setUrl] = React.useState('')
const handleExtract = async () => {
if (!url.trim()) return
const text = await extract({ url: url.trim() })
if (text) {
console.log('Extracted content:', text)
}
}
return (
<div>
<div className="flex gap-2">
<input
value={url}
onChange={(event) => setUrl(event.target.value)}
placeholder="Enter URL to extract..."
className="flex-1"
/>
<button onClick={handleExtract} disabled={isLoading || !url.trim()}>
{isLoading ? (
<span className="mr-2 inline-block animate-spin rounded-full border-2 border-current border-t-transparent" />
) : null}
Extract
</button>
</div>
{error ? <p className="text-destructive">{error.message}</p> : null}
{extractedText ? (
<div className="rounded-lg border p-4">
<Markdown>{extractedText}</Markdown>
</div>
) : null}
</div>
)
}
AI Processing
Combine extraction with AI text hooks when you want to summarize or analyze a page.
import * as React from "react";
import { useAIText } from "@/hooks/use-ai";
import { useWebExtract } from "@/hooks/use-web-extract";
export default function App() {
const { extract, isLoading: isExtracting } = useWebExtract();
const { generateText, textStream, isLoading: isSummarizing } = useAIText({
systemPrompt: "You are a helpful assistant that summarizes articles.",
});
const [url, setUrl] = React.useState("");
const handleSummarize = async () => {
if (!url.trim()) return;
const content = await extract({ url: url.trim() });
if (content) {
await generateText({
prompt: `Please summarize this article:\n\n${content}`,
});
}
};
const isLoading = isExtracting || isSummarizing;
return (/* render URL input, loading state, and textStream */);
}
Best Practices
- Pass fully qualified URLs and let the hook validate URL format before requesting extraction.
- Extract one page at a time per hook instance. Use multiple instances for concurrent extractions.
- Render or process
extractedTextas markdown. - Use
setExtractedText(null)when users clear the URL or start a new unrelated workflow. - Combine with
useAITextoruseAIObjectonly after extraction succeeds. - Keep prompts concise when sending extracted content to AI so the task stays clear.
- Surface
error.messagenear the URL input so users can correct invalid or inaccessible URLs.