You publish a genuinely good piece on a topic you have worked in for eleven years. Two weeks later you ask ChatGPT the exact question the article answers. It gives a solid response, cites four sources, and none of them are you. One of the four is a thin listicle from a content farm you have never heard of.

You check the article. It is indexed. It ranks page one. The writing is better than the listicle by any measure a human would apply.

Then you click your own byline and land on a page that says “Posted by Admin,” with an empty bio field and a default gravatar. That page is the problem, and it is a bigger problem than most content teams believe.

Answer engines are trying to solve a hard verification problem: given ten pages that say roughly the same thing, which one came from someone who knows what they are talking about. They cannot read your resume. They cannot call your references. They resolve it by checking whether the person attached to the claim exists as a verifiable entity anywhere else in the world. If your byline resolves to nothing, the engine treats the content as unsourced. That is what it means to optimize author pages for AI: you are giving the machine something to check.

Why does AI skip your author page?

Remote worker typing at a small laptop in a plain workspace, the anonymous byline problem in one image

Three reasons, in descending order of how often I find them.

The page is a feed, not a page. Most content management systems generate author archives on their own, and what they produce is a list of post titles with a name at the top. There is no prose, no claims, no dates, no credentials. To a retrieval system that page has almost no extractable content about the person. It is a menu element wearing a person’s name.

Test yours in ten seconds. Copy the entire text of your author page, paste it into a plain text file, and delete every article title. What is left? On most sites the answer is a name, a job title, and two sentences of filler. That remainder is everything the machine has to work with when it decides whether to trust anything you published.

The identity does not resolve outward. Your author page says “Sarah Chen, Head of Content.” There are a great many people named Sarah Chen. Without links that tie this specific Sarah Chen to a specific LinkedIn profile, a specific conference bio, a specific published book or a specific university page, the engine has a string, not an entity. Strings do not carry authority. Entities do.

The bio contradicts itself across properties. The website says twelve years of experience, LinkedIn says nine, the conference speaker page says “over a decade,” and the podcast intro says she started in 2015. A human reads that as normal sloppiness. A model reads it as low-confidence data and discounts the whole cluster. Inconsistency is worse than absence, because absence is neutral and contradiction is a negative signal.

Underneath all three sits Google’s E-E-A-T framing, which added Experience to Expertise, Authoritativeness, and Trustworthiness in December 2022. The reason that addition matters here is that experience is the hardest of the four to fake and the easiest to state plainly. “I ran email for a DTC brand from 2018 to 2023 and sent roughly four thousand campaigns” is a checkable claim about lived work. “Passionate about marketing” is not a claim at all.

What does an author entity actually need?

Portrait of a woman against a dark backdrop, the clear identifiable headshot every author page should carry

I call the working structure the Author Entity Spine. Six fields, all on one page, all in crawlable text. Not six nice-to-haves. Six load-bearing fields, and a page missing any two of them tends to underperform.

Field one is the canonical name string. One spelling, one form, everywhere. If you publish as “Michael J. Torres” on your site, do not publish as “Mike Torres” on Substack and “M. Torres” in the journal byline. Pick the version you want the engines to consolidate around and enforce it in every profile you control. This costs nothing and fixes more than it should.

Field two is a one-sentence identity claim that names a role, an organization, and a domain of work. “Michael J. Torres is the director of clinical operations at Northside Surgical Group, where he has managed perioperative staffing since 2016.” That sentence contains four verifiable hooks. Compare it to “Michael is a healthcare leader with a passion for patient outcomes,” which contains zero.

Field three is experience stated in checkable units. Years, counts, named employers, named clients where you are allowed, named projects, named credentials. Numbers you can defend, and nothing you cannot. If you have written 340 articles about commercial insurance, say so. If you have no idea how many, do not guess a number, name the publications instead.

Field four is the credential and affiliation set. Degrees with institutions, licenses with issuing bodies, board memberships, professional association memberships, certifications with the certifying organization named in full. For academics and researchers this includes an ORCID identifier, which is a real persistent identifier registry and one of the cleanest disambiguation signals available to anyone eligible for one.

Field five is the proof-of-work list. Not your last ten blog posts. Your external footprint: bylines in publications you do not own, conference talks with the event and year, podcast appearances, books, patents, quoted expert commentary in news coverage, court testimony if that is your field. Every item should be a link out to a property somebody else controls.

Field six is the sameAs graph, which is important enough that it gets its own section.

Work the six in order and do not skip ahead. Teams that set out to optimize author pages for AI almost always start with schema markup, because it feels like the technical part and technical parts feel like progress. Schema on an empty bio describes nothing. Write the sentences first, gather the external proof second, and mark it up last.

How do you build the sameAs graph?

The sameAs property in schema.org markup is the single most direct way to tell a machine that the person on this page is the same person as the profile over there. It is an array of URLs. It is not complicated. It is also missing from the large majority of author pages I audit.

Order matters more than volume. Lead with the profiles that carry independent verification: a Wikipedia article if one exists, a Wikidata item, an ORCID record, a university faculty page, a government or licensing board registry entry, a Crunchbase profile, a Muck Rack profile for journalists, a Google Scholar page for researchers. These are properties where a third party had to accept the entry. Then add the social profiles: LinkedIn, X, GitHub, Instagram, YouTube. Social links help with consolidation but they carry less verification weight because anyone can create one in ninety seconds.

The links have to be reciprocal wherever the platform allows it. Your author page links to your LinkedIn, and your LinkedIn links back to your author page. That round trip is what turns two floating profiles into one confirmed identity. Unreciprocated links are claims. Reciprocated links are confirmations.

Wikidata deserves specific mention because it is underused by non-academics. It is the structured data layer behind Wikipedia, it is openly editable, it is machine-readable by design, and a well-formed Wikidata item for a person with real notability is one of the strongest disambiguation objects on the open web. The caveat is real: notability standards apply, and creating a self-promotional item for someone who does not meet them wastes your time and annoys editors who will delete it. But if you have been quoted in national press, written a published book, or hold a documented professional position, it is worth checking whether an item exists and whether it is accurate.

Does author schema still matter in 2026?

Yes, and for a narrower reason than most people give.

Schema does not make an engine believe you. It removes ambiguity about what you are claiming. When a Person object sits on the page with name, jobTitle, worksFor, alumniOf, knowsAbout, description, url, image, and sameAs, the parser does not have to infer any of it from prose. That is the whole benefit, and it is worth having.

The knowsAbout property is the one people leave empty and should not. It takes an array of topics or, better, an array of URLs pointing at canonical definitions of those topics. This is where you state your subject boundaries in machine terms. Three to six topics, specific rather than broad. “Perioperative staffing” beats “healthcare.” “Series A fundraising for climate hardware” beats “startups.”

Connect the Person object to your content with author markup on every article, using a URL reference back to the author page rather than a bare text string. A bare string creates a new floating entity on every post. A URL reference points every article at one consolidated identity, and consolidation is the entire game.

One more piece that gets forgotten: the image. Use a real photograph of the person, at a stable URL, with the person’s name in the file name and in the alt attribute. Illustrated avatars and stock headshots break the visual identity chain across properties, and cross-property image matching is a live signal.

The multi-author site problem

Publications with dozens of contributors face a version of this that individual consultants do not, and the usual response makes it worse.

The usual response is a single generic template applied to everyone: same bio length, same structure, same three social icons, most fields blank because nobody chased the contributors for details. What that produces is fifty author pages that all look like low-information stubs, which drags down the perceived quality of the whole publication rather than lifting any individual writer.

The better approach is triage. Identify the eight to twelve contributors who produce the content you most want cited, and build complete Author Entity Spines for those people only. Full bios, real sameAs graphs, external proof of work, correct schema. Leave the occasional guest contributors on a simple template. A dozen strong author entities on a site outperform fifty weak ones, because the engine is consolidating authority per person, not averaging it across the masthead.

And retire the house byline where you can. “Editorial Team,” “Staff,” and “Admin” are dead ends. If a piece genuinely has no single author, attribute it to the organization as an Organization entity rather than to a fake person, and let the organizational authority carry it. That is a legitimate structure. A ghost byline is not.

Rand Fishkin is the cleanest working example of this I can point to. He founded Moz, then founded SparkToro, wrote Lost and Founder, speaks under his own name, and has published under a consistent identity for close to two decades across a company blog, a personal blog, conference stages, and a book with an ISBN. Ask an engine who he is and it answers without hedging, because every source it can reach agrees with every other source. Nothing about that is a trick. He just never split his identity across four versions of his own name, and the machines rewarded him for it years before anyone thought to optimize author pages for AI on purpose.