Thirty-four pages on this site had no h1, and here is how that happens

Thirty-three landing pages on instinctor.com rendered with no heading elements at all. Not a wrong heading, not a badly worded one. Zero h1 through h6 in the entire document, except the h3 in the cookie banner. A thirty-fourth page, the features and pricing page, carried twelve h2 section titles under no h1, and because that page is translated, the same hole existed on its Dutch, German, French, Spanish, Italian, Portuguese, Polish, Swedish, Danish and Hebrew URLs too.

Every one of those pages looked correct. Big headline at the top, section titles down the page, the right sizes, the right weights. Looked at, they were fine. Read by a machine, they had no structure whatsoever.

The words were there, the structure was not

This is the part worth understanding, because it is what makes the bug invisible.

The headline text was in the HTML the whole time. The animated headline blocks on those pages ship their text inside a visually hidden span so that a crawler reading the source sees the real words before any JavaScript runs. The words were never missing. What was missing was the tag around them: the words sat inside a div instead of an h1.

So a crawler fetching the page got the complete text of every headline, with nothing to say that any of it was a heading. To a machine the page was one long undifferentiated block of sentences, some of which happened to be set in 48px.

That is why nobody caught it by looking. A screenshot of a broken page and a screenshot of a fixed page are the same screenshot.

Why a renderer ends up defaulting to div

Both places this went wrong made the same reasonable decision.

The text element renderer takes an optional setting for its semantic tag, accepts h1 through h6, p or div, and defaults to div for anything else. The animated headline block takes a tag setting too, accepts h1, h2 or h3, and falls back to div for anything else, including unset.

Those defaults are the safe choice from the renderer's point of view. A text box is not necessarily a heading. Guessing that every large piece of text is an h1 would produce pages with nine h1s, which is worse than none. Defaulting to div means a page that has not opted in renders exactly the way it rendered yesterday.

The failure is not in either default. It is that opting in was a setting somebody had to remember, on every headline, on every page, with no warning anywhere when they did not. On thirty-three pages created in one batch, nobody did, five headlines at a time: one hero line and four section titles, one hundred and sixty five headlines painted as divs. On the features and pricing page, the section titles were tagged correctly and only the hero title was not, which is the more insidious version, because the page looks structurally healthy until you notice the twelve h2s have nothing above them.

What it actually costs

Two things, and it is worth being precise rather than dramatic about either.

For search, the h1 is the strongest in-page statement of what a page is about. A page without one is not penalized or blocked, and every one of these pages was indexed and reachable. It is simply competing without saying what it is, and relying on the title tag and body copy to carry the whole job. That is a handicap, not a disaster, and nobody can honestly tell you the exact size of it.

For accessibility the cost is concrete and easy to state. People using a screen reader navigate long pages by jumping between headings, the same way you skim by looking at bold lines. On a page with no headings there is nothing to jump to. The only way through is to listen to it from the top, in order, all of it. On thirty-three live pages, that was the only way through.

How to check your own site in about a minute

Do not trust the visual. Check the markup.

In a browser, open the page, right click, choose View Page Source, and use the find box for <h1. You want exactly one hit. Zero is the bug described here. Five means five things claim to be the subject of the page, which is its own problem.

From a terminal, one page at a time:

curl -s https://yourdomain.com/some-page | grep -c '<h1'

Run it against your home page, your main service page, and any batch of pages that were created together from the same template. Batches are where this hides, because whatever was forgotten once was forgotten identically thirty-three times.

Two things to look for beyond the count. The h1 should say what the page is about in the words a customer would use, not be a slogan or the company name. And the h2s below it should read, on their own, as a summary of the page. If you list the headings and cannot tell what the page covers, the headings are decoration.

What changed here

On the thirty-three pages, the first headline on each got tagged h1 and the remaining four got h2. On the features and pricing page, the hero title got h1. That is the whole fix: one setting per headline, applied through the normal save path.

No copy was rewritten, no element was added or removed, no layout setting was touched. The extracted body text of all thirty-three pages after the change is byte-identical to the text captured before it, which is the point. The change tells a machine what the page already said to a human. You can see the result on what a website really costs, which now renders one h1 and four h2s, the same shape the rest of the site had all along.

The general lesson, which is not about h1s

A page has two outputs. One is what it looks like, which anybody can check by opening it. The other is what it says to software: the headings, the alt text, the link text, the title and description, the structured data. Nothing on the screen tells you the second one is wrong.

So the second output needs its own check, and the check has to be mechanical, because human review of a page that looks right will pass it every time. Thirty-four pages here looked right the whole time they were wrong.

If pages on your own site are not being found, missing headings are worth ruling out early, though they are rarely the whole story. The more common reasons are covered separately in why your pages are not in Google yet.

Instinctor