LLM optimization (LLMO): be in what the model already knows, not only in what it looks up
Your company, products, and definitions present in the corpora large language models are trained on and retrieve from, so that assistants answer about you correctly even when they do not search, measured across model versions rather than months.
For companies in categories where buyers ask an assistant for a shortlist and the assistant answers from memory, and for anyone who has found that a new model version knows their competitors and not them.
Updated
What is LLM optimization (LLMO)?
LLM optimization (LLMO), sometimes called LLM SEO, is the work of making a company and its facts present in the material large language models learn from during training and retrieve from at answer time, so that assistants such as ChatGPT, Claude, Gemini, and Copilot answer about the company correctly even when they do not run a live search. It differs from answer engine optimization and generative engine optimization in horizon and target. AEO and GEO target retrieval: a page fetched today and cited today. LLMO targets the model itself: being described consistently and often enough, on sources that end up in training corpora and retrieval indexes, that the next model version knows you. In practice that means publishing reference content other sites copy and cite, making your documentation and definitions openly available and licensed for reuse, appearing in the datasets and reference works models are built from, offering a machine-readable summary of your site through llms.txt, and measuring what each new model version says before and after. At Seorythm LLMO is the long-horizon layer of the AI search programme, and it is honest about its timescale: results appear on model release cycles, not monthly ones.
What you get
| Deliverable | Cadence | What it includes |
|---|---|---|
| Model knowledge audit | Weeks 1 to 2 | What each major model states about your company and category with browsing disabled, versus with it enabled. Which of your facts the model holds, which it hedges on, which it gets wrong, and which competitors it names unprompted. Repeated for the current and previous model versions where available. |
| Reference content plan | Month 1 | The three to six definitive pieces your category lacks (definitions, comparisons, methods, data) that other sites would cite and models would learn from, with stable URLs and an owner for keeping each current. |
| Reference content production | One piece per month | Written by a named person, reviewed by your expert, with original data or a documented method, a plain summary at the top, and a licence that permits quotation and reuse with attribution. |
| Machine-readable site summary | Month 1 | llms.txt and llms-full.txt describing the site and its key pages in plain text, kept in sync with the canonical fact set, plus crawler access rules that admit the AI user agents you choose to admit. |
| Corpus-level presence | Monthly target set in roadmap | Mentions and citations of the reference content on the sources that feed training and retrieval (reference sites, documentation hubs, industry publications, high-traffic forums, open datasets where relevant). Handled with our link building team. |
| Version-over-version tracking | Monthly checks, reported at each model release | The audit prompt set re-run against each new model version with browsing off and on, fact accuracy and unprompted mention rate recorded, and the change attributed to retrieval or to the model where the evidence allows. |
How the work is done
Ask with browsing off
A model answers in two ways: from what it holds, and from what it fetches. The audit separates them. We ask each major model about your company and category with browsing disabled, then again with it enabled, several times each, and record the difference. The first set shows what the model learned; the second shows what it can find. Where a fact is right only with browsing on, retrieval is carrying you and the model does not know you. Where a competitor is named unprompted with browsing off, the model learned them. That gap, per fact and per competitor, is what LLMO works on, and the AI optimization fact set is its starting point.
Write the reference your category lacks
Models learn what is written often, copied widely, and cited by others. In most categories there are three to six pieces that do not yet exist in definitive form: the plain definition, the honest comparison, the method explained step by step, the dataset nobody has published. The plan names them. Each is then written by a named person, reviewed by your expert, and given the things that make a page get quoted: a one-paragraph summary at the top, original data or a documented method, specific figures with their source, and a stable URL that will not change at the next redesign. One piece a month is the usual pace; a reference page that is wrong or thin does more harm than none.
Make it stable, open, and copied
A page that is going to be learned from has to be reachable, permanent, and reusable. We publish reference content under a licence that permits quotation and reuse with attribution, keep it at the same URL for years, and then work to get it cited and quoted on the sources that feed training and retrieval: reference sites, documentation hubs, industry publications, the high-traffic forums in your category, and open datasets where the content fits. That is link building with a different scorecard, in which a verbatim quote with your name and no link still counts. We do not flood forums, seed promotional text, or attempt to poison anything; the providers filter for exactly that, and a company caught doing it loses the legitimate presence too.
Tell the machines what the site is
We publish llms.txt and llms-full.txt, plain-text summaries of the site and its key pages kept in step with the canonical fact set, and set crawler rules for the AI user agents you choose to admit. llms.txt is a proposal, not a standard, and we will not tell you a provider has committed to it; it is cheap, harmless, and read by some retrieval systems, which is enough to justify an hour. Crawler policy is a real decision with a real trade-off, and it is yours: blocking keeps content out of training and out of retrieval-based answers at the same time. This is technical SEO work with the user-agent list changed.
Report on the model’s clock
LLMO is the one layer of our programme that does not move monthly, and the reporting says so. We re-run the audit prompts every month so that retrieval changes are caught, but the headline numbers, fact accuracy with browsing off and unprompted mention rate for a fixed set of category prompts, are reported per model version, because that is when they can change. Each is sampled and reported as a rate with a noise band. When a new version knows you and the previous one did not, you will see it; when it does not, you will see that too, next to the answer engine and generative engine citation numbers that the same content earned in the meantime.
LLMO is the long-horizon layer of the AI search programme. Where it sits next to the others is set out in SEO vs AEO vs GEO vs AIO vs LLMO vs SXO. To find out what the models know about you today, start with the audit, or see pricing.
What it costs
Included in the AI search programme, or as a standalone LLMO project
The model knowledge audit is part of the $2,800 audit. Ongoing LLMO work runs inside the sprint or retainer, or within the standalone AI search programme at $1,200 to $4,000 a month alongside AEO, GEO, and AIO. Reference content production is priced per piece within those programmes.
What moves the number
- Number of reference pieces to produce and how much original data each needs
- How thin your category's existing reference material is
- Number of models and versions tracked
- Whether your facts already agree across sources or need AIO work first
Who this is not for
- Anyone who needs results this quarter. LLMO moves on model release cycles. If you need citations now, buy the AEO and GEO work; it uses the same content and retrieves within weeks.
- Companies unwilling to publish anything openly. Content behind a login or a paywall is not learned from and rarely retrieved.
- Buyers who want us to poison, spam, or flood training sources with promotional text. It does not work, it is detectable, and the model providers are removing exactly that material.
- Companies with no distinct facts, data, or method to publish. Reference content has to be worth referencing.
Risks and how we manage them
| Risk | How we manage it |
|---|---|
| Money spent, and the next model version still does not know you. | The retrieval half of the work pays back within weeks through AEO and GEO citations, so the programme is not a bet on a single release. The version-over-version report shows unprompted mention rate honestly, including when it has not moved. |
| A model provider changes its crawling or training policy. | We build presence on the sources every provider uses (reference sites, documentation, publications, forums) rather than on any single crawler, and we keep the crawler rules under your control. |
| Reference content is copied without attribution. | The licence requires attribution, the content states its origin in the text itself, and a copied definition that names you is still doing its job. |
| Reading a model's random variation as movement. | Each prompt is sampled several times per model per version and we report rates, not single answers. Changes inside the noise band are labelled as such. |
Can you influence what an LLM is trained on?
Partly, indirectly, and only in ways that are legitimate. Nobody outside a model provider chooses what goes into training, and vendors who claim otherwise are selling something they do not have. What a company can do is change what exists to be trained on. Training corpora and retrieval indexes are built from the open web weighted toward pages that are widely linked, frequently copied, stable over time, and permissively licensed, plus reference works, documentation, forums, and news. A company that publishes the definitive plain-language explanation of its category, keeps it at a stable URL for years, licenses it for reuse, gets it cited and quoted by other sites, and states its own facts identically everywhere is far more likely to be represented in the next model than one whose content is thin, changes URL every redesign, or lives behind a login. The other lever is retrieval-augmented answering: most assistants now fetch live pages for many prompts, and the same content that trains well also retrieves well. LLMO works both, and reports them separately.
Questions buyers ask
Is LLMO a real discipline or a rebrand of SEO?
It is a real, narrow extension. Most of it is content people should have been producing anyway: definitive, stable, openly licensed, widely cited reference material. What is new is the target (a model's knowledge, not a page's rank), the time horizon (release cycles), and the measurement (asking the model with browsing off, across versions). Anyone selling it as unrelated to SEO, or as fast, is overselling.
What is the difference between LLMO and AEO or GEO?
AEO and GEO target what an engine retrieves and cites today; they move in weeks. LLMO targets what a model already holds; it moves when models are retrained. The content overlaps almost completely, so we sell all of them as one AI search programme and report retrieval and model knowledge separately.
How much does LLM optimization cost?
The model knowledge audit is included in our $2,800 audit. Ongoing work runs inside the sprint, the retainer, or the standalone AI search programme at $1,200 to $4,000 a month; reference content is priced per piece within those. The comparison is on the pricing page.
Does llms.txt actually help?
It is a proposed convention, not a standard any provider has committed to, so treat it as cheap and possibly useful rather than proven. It costs an hour, keeps a plain-text summary of your site in one place, and some retrieval systems read it. We publish it and do not overclaim for it.
Should we block AI crawlers?
That is your decision and it is a trade-off. Blocking protects content from being used in training and removes you from retrieval-based answers at the same time. We show you which user agents do what, set the rules you choose, and revisit them when providers change their behaviour.
How do you measure LLMO?
Two rates per model version, sampled. Fact accuracy with browsing disabled, which reflects what the model holds. Unprompted mention rate for a fixed set of category prompts, which reflects whether the model thinks of you at all. Both are reported alongside the retrieval-based citation numbers from the AEO and GEO work, so you can see which layer is doing what.
Start with the audit.
Two weeks, fixed price, and a prioritised list you can act on with or without us. If SEO is not the right channel for you, the audit will say so.