Content Signals in robots.txt: let AI cite you without training on you
Content Signals in robots.txt
Content Signals are a line you add to robots.txt that says what AI systems may do with your pages after they crawl them. There are three signals: search, ai-input and ai-train, each set to yes or no. To let AI tools read and cite you without training on you, add Content-Signal: search=yes, ai-input=yes, ai-train=no under your User-agent group. That is the exact line linkinseconds.com publishes.
Below: what each signal means, why we chose our values (user-uploaded files were the deciding factor), how to add the line in Next.js, and what Content Signals can and cannot do.
The gap robots.txt leaves open
Classic robots.txt answers one question: may this crawler fetch this path? It has nothing to say about what happens next. A page fetched to build a search index, a page fetched to answer one user's question in a chatbot, and a page fetched to train the next model all look the same to robots.txt. You can block a crawler or allow it. You cannot say “yes to this use, no to that one”.
That is a real problem, because many sites want search traffic and AI citations but do not want their writing folded into training data. Blocking AI crawlers outright loses the first two. Allowing them says nothing about the third.
What the three signals mean
Content Signals were proposed by Cloudflare and are defined at contentsignals.org. In short:
| Signal | Covers | Our value |
|---|---|---|
| search | Building a search index and showing search results: links and short excerpts | yes |
| ai-input | Feeding the page into an AI model at answer time, such as retrieval, grounding or AI search answers | yes |
| ai-train | Training or fine-tuning AI models | no |
Note that search does not include AI-written summaries. Those fall under ai-input. If you leave a signal out, you express no preference about that use either way.
Our line, and why
User-Agent: *
Allow: /
Disallow: /dashboard
Disallow: /api/
Disallow: /login
Content-Signal: search=yes, ai-input=yes, ai-train=no
Sitemap: https://linkinseconds.com/sitemap.xml- search=yes. We want people to find our guides. That one is easy.
- ai-input=yes. When someone asks an AI assistant “how do I send a file too big for email?”, we would like it to read our answer and link to it. That is the modern version of a search result.
- ai-train=no. This is the important one, and the reason is our users.
User-uploaded files changed the answer
Link in Seconds turns files into public links. People share resumes, invoices, pitch decks, photos and whole websites. Those pages live on our domain, but the content belongs to the people who uploaded it, not to us. We have no right to offer it up for model training, and our users would not expect us to.
Our Content-Signal line sits in the group that covers every path, so ai-train=no applies to file pages under /p/ and album pages under /b/ as well as to our own guides. Those file pages also carry noindex, so they stay out of search results. If your site hosts anything your users made, think about their content first and your marketing pages second.
How to add Content Signals in Next.js
If your robots.txt is generated by app/robots.ts, you do not need to switch to a static file. Recent versions of Next.js support an other field on each rule. Every key in it is printed as a line inside that rule's User-Agent group:
// app/robots.ts
import type { MetadataRoute } from "next";
export default function robots(): MetadataRoute.Robots {
return {
rules: {
userAgent: "*",
allow: "/",
disallow: ["/dashboard", "/api/", "/login"],
// Printed as "Content-Signal: ..." inside this User-Agent group
other: { "Content-Signal": "search=yes, ai-input=yes, ai-train=no" },
},
sitemap: "https://example.com/sitemap.xml",
};
}We keep the value in one exported constant, so our docs and this post quote the same string the robots file serves. If you use a static public/robots.txt, just add the line under User-agent: * by hand.
Check that it is live:
curl -s https://linkinseconds.com/robots.txt | grep -i content-signalMistakes to avoid
- Putting the line outside a group. Content-Signal belongs inside a
User-agentblock, like Allow and Disallow. A line floating at the bottom of the file next toSitemapmay be ignored. - Saying no to ai-input when you want AI citations. If you want assistants to quote and link you, they need to be allowed to read you at answer time.
- Blocking AI crawlers and adding signals. If a bot is disallowed, it never reads your signals. Pick one approach per crawler.
- Forgetting user content. Your values apply to every path in the group, including pages your users created.
Does it change anything measurable?
On Cloudflare's isitagentready.com scanner, Content Signals is one of the Bot Access Control checks. Before we added the line we scored 1 out of 2 there. After, 2 out of 2. Together with a small routing fix and Markdown for agents, it moved our overall score from 27 to 40. The full write-up is in Is your site agent-ready?
Common questions
Is Content Signals an official standard?
It is a published proposal with a public spec at contentsignals.org, not an IETF standard. It is designed to sit inside the robots.txt format that crawlers already read.
Do I need llms.txt as well?
They do different jobs. Content Signals say what AI may do with your pages. An llms.txt file tells an AI what your site is and which pages matter. We publish both.
Where can I read your full crawler policy?
In AI crawlers, Markdown and content signals. For the files that describe our API and MCP server, see Agent discovery files, and for more guides like this, the AI agents and developers topic.
Turn any file into a link in seconds
Upload a PDF, image, video, or ZIP and get a clean, trackable link with a QR code, free.
Try Link in Seconds →