All articles

Serve Markdown to AI agents with content negotiation (Next.js)

Sunny 8 min read

#ai#developers#markdown#agents#nextjs

Serve Markdown to AI agents

pricing.md Live
/p/pricing

To serve Markdown to AI agents in Next.js, use content negotiation: when a request sends Accept: text/markdown, rewrite it in middleware to a route handler that returns a Markdown version of the same page, with Content-Type: text/markdown and Vary: Accept. Browsers never ask for Markdown, so they keep getting HTML at the same URL. Agents get clean text that costs far fewer tokens.

This is exactly how linkinseconds.com works today. Try it:

Shell
curl -s -H "Accept: text/markdown" https://linkinseconds.com/pricing

Below are the pieces, simplified from our real code: the Accept check, the middleware rewrite, the route handler, the converter, and the headers that keep caches and search engines from mixing the two versions up.

Why Markdown, and why content negotiation?

An agent that fetches a normal web page gets navigation, a footer, scripts, inline SVG icons and a lot of class names. It has to strip all of that to find the few paragraphs that matter. Markdown is the opposite: headings, lists, links and text, nothing else. It is the format language models read most easily.

You could publish a separate /page.md for every page, but then agents have to know that convention. Content negotiation is the web's built-in answer: one URL, several representations, and the client picks with its Accept header. No new URLs, nothing to keep in sync, and nothing changes for people.

Step 1: read the Accept header properly

Do not just check whether the header contains the string text/markdown. Accept headers carry quality values (q=0.8), and a client might list Markdown as a last resort behind HTML. We only switch when Markdown is preferred at least as much as HTML. A wildcard */* does not count as asking for HTML, so an agent sending text/markdown, */*;q=0.8 gets Markdown.

markdown-negotiation.ts
// Does this Accept header ask for Markdown at least as strongly as HTML?
export function prefersMarkdown(accept: string | null): boolean {
  if (!accept) return false;
  const q = new Map<string, number>();
  for (const part of accept.split(",")) {
    const [type, ...params] = part.split(";");
    let weight = 1;
    for (const p of params) {
      const [k, v] = p.split("=");
      if (k?.trim() === "q") weight = Number(v) || 0;
    }
    q.set(type.trim().toLowerCase(), weight);
  }
  const md = q.get("text/markdown") ?? 0;
  if (md <= 0) return false;
  const html = Math.max(q.get("text/html") ?? 0, q.get("application/xhtml+xml") ?? 0);
  return md >= html; // a bare */* never counts as asking for HTML
}

Step 2: rewrite in middleware

A rewrite, not a redirect. The visible URL stays the page's own, which is what content negotiation promises. In our app this rule runs before locale routing and before the “signed-in visitors go to the dashboard” redirect, so an agent always gets the public page.

middleware.ts
import { NextResponse, type NextRequest } from "next/server";
import { prefersMarkdown } from "@/lib/markdown-negotiation";

const SELF_FETCH = "x-markdown-fetch";

export function middleware(request: NextRequest) {
  if (
    request.method === "GET" &&
    !request.headers.has(SELF_FETCH) &&
    prefersMarkdown(request.headers.get("accept"))
  ) {
    const url = request.nextUrl.clone();
    const page = url.pathname;
    // Put the page path in the pathname: a rewrite's query string
    // does not reach the route handler.
    url.pathname = page === "/" ? "/agent-markdown" : `/agent-markdown${page}`;
    url.search = "";
    return NextResponse.rewrite(url);
  }
  // ...your normal middleware (i18n, auth) continues here
}

Two details that cost us time. First, a rewrite's query string does not reach the route handler, so the page path rides in the pathname (/agent-markdown/pricing) and is read from a catch-all segment. Second, the route fetches the page's own HTML, so that fetch carries a marker header and the middleware skips it. Without that, the route would rewrite to itself forever.

On Next.js 16 you will see a warning that middleware is becoming proxy. The rewrite works the same either way.

Step 3: the route handler

app/agent-markdown/[[...path]]/route.ts
// app/agent-markdown/[[...path]]/route.ts
import sitemap from "@/app/sitemap";
import { htmlToMarkdown } from "@/lib/markdown-negotiation";

const allowed = new Set(sitemap().map((e) => new URL(e.url).pathname));

function markdown(body: string, status = 200) {
  return new Response(body, {
    status,
    headers: {
      "Content-Type": "text/markdown; charset=utf-8",
      "x-markdown-tokens": String(Math.ceil(body.length / 4)),
      Vary: "Accept",
      "Cache-Control":
        status === 200 ? "public, max-age=0, s-maxage=3600, stale-while-revalidate=86400" : "no-store",
      "X-Robots-Tag": "noindex",
    },
  });
}

export async function GET(req: Request, ctx: { params: Promise<{ path?: string[] }> }) {
  const { path: parts } = await ctx.params;
  const path = "/" + (parts ?? []).join("/");
  if (!allowed.has(path)) {
    return new Response("No Markdown version of this page.\n", {
      status: 406,
      headers: { Vary: "Accept", "Cache-Control": "no-store" },
    });
  }

  // Fetch our own HTML, signed out (no cookies forwarded), and convert it.
  const origin = new URL(req.url).origin;
  const res = await fetch(new URL(path, origin), {
    headers: { accept: "text/html", "x-markdown-fetch": "1" },
    cache: "no-store",
    signal: AbortSignal.timeout(10_000),
  });
  if (!res.ok) return markdown("# Not available\n", 502);
  return markdown(htmlToMarkdown(await res.text(), origin));
}

Our English home page is a special case: it answers with our llms.txt, which is already a hand-written Markdown summary of the whole site. That is a better answer for “what is this site?” than a converted marketing page.

Step 4: a small, safe converter

We did not pull in a general HTML to Markdown library. We only ever convert our own server-rendered pages, so the converter can be small and strict:

  • It reads only what is inside <main>. Header, footer and cookie banners stay out.
  • It drops scripts, styles, SVG, iframes, buttons and forms entirely.
  • It maps headings, paragraphs, lists, links, bold, italics, inline code and tables.
  • Relative links become absolute, so an agent can follow them.
  • It is a single forward pass with no backtracking patterns, so it stays fast even on malformed HTML. A regex that can backtrack is a denial of service bug waiting to happen.

It is unit-tested on its own, which is easy because it is a pure function: HTML in, Markdown out.

The headers that matter

  • Content-Type: text/markdown; charset=utf-8. Tells the client what it got.
  • Vary: Accept. The most important one. It tells every cache (your CDN, a proxy, the browser) that this URL has more than one version depending on Accept. Leave it out and a CDN can cache the Markdown and hand it to the next person who opens the page in a browser.
  • x-markdown-tokens. A rough token count (characters divided by four) so an agent can decide whether the page fits its budget before reading it.
  • X-Robots-Tag: noindex. Search engines should index the HTML page, not a second copy of it.
  • Cache-Control. We cache at the edge for an hour and serve stale while revalidating. Errors are no-store.

Why only indexable pages

This is the part to get right. The route fetches pages and returns their content, so it must never become a way to read something private. Our rule: only paths in the sitemap (and their translated versions) get a Markdown version. Everything else, including the dashboard, sign-in pages and every user's uploaded file, gets a 406 Not Acceptable and no content.

Using the sitemap as the allowlist means there is one list of public pages, not two. The self-fetch also sends no cookies, so even an allowed page is rendered exactly as a signed-out visitor would see it.

Markdown negotiation is one of the checks on Cloudflare's isitagentready.com scanner. Adding it moved our Content score from 0/1 to 1/1. The full story is in Is your site agent-ready?

A checklist for your own site

  1. Parse Accept with q-values. Switch only when Markdown wins.
  2. Rewrite, do not redirect. Keep the URL.
  3. Mark the self-fetch so the middleware cannot loop.
  4. Allowlist public pages. Return 406 for everything else.
  5. Send Vary: Accept on every response from the route, errors included.
  6. Add noindex to the Markdown version.
  7. Test with curl, then open the page in a browser to confirm nothing changed for people.

Common questions

Will this hurt my SEO?

No. Search engine crawlers ask for HTML, so they see what they always saw. The Markdown version carries noindex in case one ever lands on it.

Do AI crawlers actually send Accept: text/markdown?

Some do: agent-readiness scanners such as isitagentready.com send exactly this header to test a site. Many crawlers still ask for HTML, which is why we also publish llms.txt and state our preferences with Content Signals in robots.txt.

Where is this documented?

See AI crawlers, Markdown and content signals and Agent discovery files, or browse the AI agents and developers guides.

Turn any file into a link in seconds

Upload a PDF, image, video, or ZIP and get a clean, trackable link with a QR code, free.

Try Link in Seconds →