SEO for News and Content Publishers
Publisher problems are rarely about individual articles. They are about a library that grew for a decade without anyone deciding what should still be in it, crawl attention going to the wrong years, and a reputation Google judges all at once. I have worked inside a 100M pageview operation and recovered a publisher from a Google manual action.
Publishing operations I have worked inside
Where publisher traffic quietly leaks
None of these appear in a report covering the last thirty days, which is how they survive for years. All of them are structural, and every one can be fixed without publishing anything new.
The archive nobody has opened since 2019
Every publisher has hundreds of thousands of pages about things nobody will search for again. They are not harmless. They absorb crawl attention, blur what the site looks like it is about, and make the pages that still matter harder to reach.
Tag and category pages the CMS invented
Most publishing platforms generate a page for every tag anyone has ever typed. The result is thousands of near empty listing pages, indexable by default, each one a candidate for crawl attention and none of them useful to a reader.
Old articles competing with new ones
You covered the same recurring event five years running and now have five pages fighting over one query. Google picks one, often not the one you would have chosen, and the traffic splits until somebody decides which page is the page.
Pagination that goes nowhere
Section archives running to page 400 with no other route into the content underneath. Crawlers walk a few pages deep and stop, which leaves a large part of the site technically linked and practically unreachable.
Syndicated content with no canonical
Wire copy and partner content republished without a canonical pointing home. At volume this teaches Google that a meaningful share of the site is other people's writing, and that is a judgement about the publication rather than the page.
Everything depending on one surface
Traffic that is 70% Discover is not a strategy, it is exposure. One core update and the quarter is gone. Knowing the split and deliberately building the weaker side is a Discover question and a Search one at once.
What I work on
Archive strategy
Deciding what gets kept, merged, archived or removed across a library nobody has audited in years. Unglamorous, slow, and usually where a large publisher's crawl efficiency went. Also the work most reliably postponed.
Crawl budget at publisher scale
Log file analysis to see where Googlebot actually spends its time across templates and decades of content, then changing the structure so it spends it better. News sitemaps help new stories get found; they do not fix a site where the crawl is being spent elsewhere. This is the technical work applied at volume.
Recovering publisher trust
I recovered a sports media site from a Google manual action at Better Collective. The part that generalises is that Google judges a publication as a publication, so the fix is never only the pages that triggered it.
Migrations without losing the archive
Publishers carry more URLs than anyone and lose more of them in a bad migration. Redirect mapping across a large archive, indexation monitored through launch, and someone watching the numbers settle for weeks afterwards.
Planning around demand
Events, seasons and fixtures create demand on a schedule. Aligning editorial output with that calendar is most of the job at a sports publisher, and the same shape applies to any predictable news cycle.
The workflow underneath it
None of this holds if the desk cannot apply it while publishing at speed. Structural decisions have to reach the brief and the template, which is my editorial systems work.
How I approach a publisher
Five things, and the fourth is the one everybody postpones.
Look at the whole library, not the last month
Publisher problems accumulate over years and are invisible in a report covering thirty days. The first thing I want is the shape of the entire site by age, template and traffic, because the answer is usually sitting in that view.
Read the logs before the tools
On a site with a million URLs, a crawler tells you what exists and a log file tells you what Google cares about. Those two lists are always further apart than anyone expects, and the gap is the work.
Treat trust as a property of the publication
Google forms a view of a publication, not just its pages. That is why a manual action is never fixed by cleaning up only the pages that caused it, and why core updates move whole sites rather than sections.
Decide what dies
Somebody has to say which of the four hundred thousand pages are not worth keeping. It is a genuinely uncomfortable conversation to have with an editorial team, and it is usually the highest value item on the list.
Reduce single surface dependence
A publisher earning most of its traffic from one place is one update away from a very bad quarter. Knowing the split and deliberately building the weaker side is strategy rather than optimisation, and it needs an owner.
"Google forms a view of a publication, not just its pages. That is why a manual action is never fixed by cleaning up only the pages that caused it."
Publishing work I can point at
Three examples with the company named. The full write ups are on the homepage.
Recovering a publisher from a Google manual action
Worked out what triggered it, cleaned it up, then rebuilt the architecture so the same thing could not accumulate again.
SEO inside a 100M pageview publishing operation
2,000+ articles optimised, over 1M organic sessions a month, and a new section that reached 300K+ keywords.
Migrations on high volume content platforms
URL structure, redirect mapping and indexation monitoring through launch and for weeks after it.
Frequently Asked Questions
Your archive is bigger than your traffic suggests
Tell me roughly how many URLs you have and where the traffic comes from, and I will tell you where I would look first, whether or not we end up working together.



