CI/CD & ARCHITECTURE · · 8 MIN READ

How Firecrawl, Docusaurus & Next.js 15 Generate llms.txt at Scale

Manually writing an /llms.txt file is manageable for a 5-page landing site. But for rapidly changing SaaS documentation and content platforms, manual curation breaks within days. Here is the production engineering blueprint to automate /llms.txt and /llms-full.txt across your CI/CD pipelines.

Architectural Reality Check

"Our documentation changes 20 times a week. When our llms.txt got out of sync with our API, ChatGPT Search started citing deprecated v1 endpoints. Automating llms.txt during build time solved hallucinated citations completely."

— Sourced from engineering discussions on Reddit r/SEO

01 The Two-Tier Protocol: llms.txt vs llms-full.txt

Before writing build scripts, understand the two-tier architectural standard adopted by LLM crawlers (GPTBot, PerplexityBot, ClaudeBot):

1. /llms.txt (Table of Contents)

Curated index of markdown links with single-line summaries. Must remain strictly under 30KB to prevent RAG context overflow during sub-second real-time search queries.

2. /llms-full.txt (Full Technical Corpus)

Uncompressed concatenation of all public documentation, API schemas, and code recipes used for deep agentic reasoning and offline index ingestion.

For a complete breakdown of why static files fail when misconfigured, see our audit: Why 97% of llms.txt Files Fail (And the 3 Indexing Traps).

02 Pipeline 1: Automated Firecrawl Web Ingestion

If your website uses a complex CMS, WordPress, or Webflow where you cannot easily access raw markdown files at build time, you can use the Firecrawl API in a GitHub Action to crawl and format your domain:

// scripts/generate-llms-firecrawl.js
import fs from 'fs';

const API_KEY = process.env.FIRECRAWL_API_KEY;
const DOMAIN = 'https://yourdomain.com';

async function generate() {
  console.log('Crawling site with Firecrawl for LLM indexing...');
  const response = await fetch('https://api.firecrawl.dev/v1/crawl', {
    method: 'POST',
    headers: {
      'Authorization': `Bearer ${API_KEY}`,
      'Content-Type': 'application/json'
    },
    body: JSON.stringify({
      url: DOMAIN,
      limit: 50,
      scrapeOptions: { formats: ['markdown'] }
    })
  });

  const { data } = await response.json();
  
  // 1. Build /llms.txt index
  let indexContent = `# Knowledge Base\n> Curated AI Index for ${DOMAIN}\n\n## Documentation\n`;
  let fullContent = `# Complete Corpus for ${DOMAIN}\n\n`;

  data.forEach(page => {
    const title = page.metadata?.title || page.url;
    const desc = page.metadata?.description || 'Technical guide';
    indexContent += `- [${title}](${page.url}): ${desc}\n`;
    fullContent += `\n---\n# ${title}\nURL: ${page.url}\n\n${page.markdown}\n`;
  });

  fs.writeFileSync('./public/llms.txt', indexContent);
  fs.writeFileSync('./public/llms-full.txt', fullContent);
  console.log('Successfully generated public/llms.txt and public/llms-full.txt');
}

generate();

03 Pipeline 2: Docusaurus Post-Build Plugin

Docusaurus stores docs in raw markdown. You can generate clean /llms.txt files automatically via a lightweight local Node script executed after docusaurus build:

// scripts/docusaurus-llms-gen.js
const fs = require('fs');
const path = require('path');
const matter = require('gray-matter');

const docsDir = path.join(__dirname, '../docs');
const outDir = path.join(__dirname, '../build');

function scanDocs(dir) {
  let results = [];
  const list = fs.readdirSync(dir);
  list.forEach(file => {
    const filePath = path.join(dir, file);
    const stat = fs.statSync(filePath);
    if (stat && stat.isDirectory()) results = results.concat(scanDocs(filePath));
    else if (file.endsWith('.md') || file.endsWith('.mdx')) {
      const raw = fs.readFileSync(filePath, 'utf8');
      const { data, content } = matter(raw);
      results.push({
        title: data.title || file.replace(/\.mdx?$/, ''),
        description: data.description || data.sidebar_label || '',
        slug: data.slug || file.replace(/\.mdx?$/, ''),
        content
      });
    }
  });
  return results;
}

const docs = scanDocs(docsDir);
let index = `# Developer Documentation\n> Technical API and architecture reference.\n\n## Guides\n`;
docs.forEach(d => {
  index += `- [${d.title}](https://yourdocs.com/${d.slug}): ${d.description}\n`;
});

fs.writeFileSync(path.join(outDir, 'llms.txt'), index);
console.log('✔ Generated build/llms.txt for Docusaurus');

04 Pipeline 3: Next.js 15 Dynamic Route Handler

In Next.js 15 App Router, you can build a zero-overhead static route handler at app/llms.txt/route.ts that compiles markdown index files during next build:

// app/llms.txt/route.ts
import { getAllArticles } from '@/lib/articles';

export const dynamic = 'force-static';
export const revalidate = 86400; // 24 hours

export async function GET() {
  const articles = await getAllArticles();
  
  let body = `# Acme Platform Index\n> Authoritative developer knowledge base.\n\n## Core Resources\n`;
  
  articles.forEach(art => {
    body += `- [${art.title}](${art.canonicalUrl}): ${art.summary}\n`;
  });
  
  body += `\n- [/llms-full.txt](https://acme.com/llms-full.txt): Complete uncompressed API reference.\n`;

  return new Response(body, {
    status: 200,
    headers: {
      'Content-Type': 'text/plain; charset=utf-8',
      'Cache-Control': 'public, max-age=86400, s-maxage=86400',
    },
  });
}

Validate Your Automated llms.txt Output

Test your domain for MIME-type compliance, crawler robots.txt clearance, and GEO citation readiness in one click.