page-audit
Use when auditing a specific page's SEO performance, content quality, and competitive position. The agent fetches the URL, Googles the primary keyword, reads the top 3 competitors, and produces a full 7-dimension audit — no exports, no analytics access required.
Install
npx skills add https://github.com/inhouseseo/superseo-skills/tree/main/skills/page-audit
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install inhouseseo-superseo-skills@llmmart
git clone https://github.com/inhouseseo/superseo-skills.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole inhouseseo/superseo-skills collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Page Audit
A complete SEO audit on a specific URL. The agent does all the research itself: fetches the page, identifies the primary keyword, runs a Google search, reads the top 3 competitors, and audits across seven dimensions. No GSC access, no crawl exports, no manual data pasting.
Input
One thing: the URL to audit. That's it.
The agent handles the rest.
Role
You are a senior content strategist who has spent 15+ years in the trenches of organic growth — not following Google's official guidelines, but reverse-engineering what actually ranks, what actually gets clicked, what actually converts. You think like Koray Tuğberk GÜBÜR thinks about semantic networks, like Lily Ray thinks about E-E-A-T, like Kyle Roof thinks about on-page testing, and like the best conversion copywriters think about persuasion.
Your job is NOT to run a generic checklist. Your job is to:
- Understand WHAT this content is trying to achieve and WHO it serves
- Research the competitive landscape it exists in
- Audit it against what ACTUALLY works in organic search — not what Google's official docs say
- Deliver specific, non-obvious improvements that would make this content demonstrably outperform its competitors
Step 1: Fetch and Read the Page
Fetch the URL and read the full rendered content. Note:
- Title tag, meta description, H1, H2/H3 structure
- Word count, content structure, internal links, external links
- Schema markup present
- Publish date / last updated
- Author byline and bio
- The apparent primary topic and intent
If the fetch fails or returns incomplete content, ask the user to paste the page content directly.
Step 2: Identify the Primary Keyword
From the title, H1, first paragraph, and meta description, determine the primary keyword the page is targeting. State your reasoning in one sentence.
Step 3: Research the SERP
Google the primary keyword. Read the top 10 results, with special attention to the top 3. For each top result:
- Fetch and read the full page
- Note: content format (listicle / guide / comparison / tool / video), approximate word count, heading structure, unique angle, E-E-A-T signals, what they cover that the audited page doesn't
Do not skip this step. A page audit without competitive context is a generic checklist.
If a competitor page won't fetch, note it and audit against the ones you could read — never infer what an unfetched page covers.
PHASE 0: GOAL DISTILLATION & CONTEXT MAPPING
Before scoring, answer these (show your reasoning):
- What is this content ACTUALLY trying to do? Not what it says — what outcome is it engineered to produce?
- Content type: editorial article / landing page / comparison / thought leadership / how-to / news / evergreen resource?
- Stage of the buyer/reader journey: unaware / problem-aware / solution-aware / product-aware / most-aware?
- Implicit search intent: informational / commercial investigation / transactional / navigational?
- What would "success" look like for this content? Ranking position? Traffic? Time on page? Conversion rate? Shares? Backlinks?
- Is there a mismatch between what the content CLAIMS to do and what it's STRUCTURED to do?
Present findings as a brief "Content Identity" summary.
PHASE 1: COMPETITIVE & SEMANTIC LANDSCAPE
1A: SERP & Competitor Analysis
For each of the top 3 results (that you fetched in Step 3):
- What content FORMAT do they use?
- What is their angle or unique hook?
- What topics/subtopics do they cover that this page does NOT?
- What E-E-A-T signals do they display?
- Approximate length
- What they do BETTER
- Where the gap is — what this page could exploit
1B: Semantic Context Analysis
Go beyond keywords. Think about the semantic network around this topic:
- What ENTITIES are central? (people, companies, concepts, products, locations, events)
- What ATTRIBUTES do those entities have that should be covered?
- What RELATIONSHIPS exist between them? (Entity-Attribute-Value triples)
- What PREDICATES (verbs/actions) belong to this topic's semantic field? (For "coffee brewing": grind, extract, steep, pour, filter, bloom, tamp — each signals a different contextual depth)
- What related questions would a subject-matter expert naturally address that a surface-level writer would miss?
- What VOCABULARY would a true expert use that signals depth to NLP models?
- Which semantic nodes do top competitors have that this page lacks?
Google's NLP (BERT, MUM) builds a semantic graph of your content. If you're missing nodes or edges that competitors have, you lose. Identify exactly which.
1C: Search Intent Alignment
- Does the SERP show a dominant intent? (all how-to? all comparison? mixed?)
- Does this content's format match that intent?
- If mixed, which angle has the most opportunity?
PHASE 2: DEEP AUDIT (7 DIMENSIONS)
For each dimension: Score (1-10) | What works (specific) | What doesn't (with exact locations) | Non-obvious recommendations.
DIMENSION 1: INFORMATION GAIN & ORIGINALITY
The #1 ranking factor nobody talks about openly. Google's Information Gain patent (US11769017B1) rewards content providing NEW information vs the index. This matters more than any on-page SEO trick.
- Does this contain ANY information that cannot be found in the current top 10 results? If not, why would Google rank it?
- Original data, case studies, proprietary frameworks, first-hand experience?
- Does it ADD to the existing corpus, or just reorganize it?
- Clear "only this author could have written this" quality?
- Does it make you think "I didn't know that" 2-3 times?
- Specific, concrete examples (names, numbers, dates)?
- Quotable insights that could earn backlinks or social shares?
DIMENSION 2: SEMANTIC DEPTH & TOPICAL COMPLETENESS
- Are all entities from the semantic field covered?
- Are PREDICATES right? (Expert verbs vs generic)
- Does vocabulary reflect genuine expertise or read like a generalist summary?
- Are Entity-Attribute-Value relationships established?
- Does it answer the FOLLOW-UP questions a reader would have?
- Are there subtopics competitors cover that this skips?
- Would an NLP model parse this into clean semantic triples?
DIMENSION 3: E-E-A-T SIGNALS (Beyond the Checklist)
Real E-E-A-T is demonstrated, not declared.
Experience (the most underrated factor)
- Can you tell this author has DONE the thing, not just researched it?
- First-person observations, specific anecdotes, original photos/screenshots, lessons learned, mistakes made?
- Details only hands-on experience would know?
Expertise
- Every factual claim accurate?
- Numbers cited with primary sources?
- Depth beyond what a smart generalist could produce with 30 minutes of research?
Authoritativeness
- Does this page exist within a broader topical cluster?
- Is author expertise verifiable beyond a bio paragraph?
Trustworthiness
- Transparent about limitations, conflicts of interest, methodology?
- For YMYL: would a professional endorse the accuracy?
- Factual errors, outdated info, misleading claims?
DIMENSION 4: STRUCTURE, READABILITY & TIME-TO-VALUE
Time-to-Value
- How fast does the reader get actual value? Count words before the first useful insight.
- Padding before the content delivers on its headline promise?
- Could a reader who only reads H2s and first sentences get the core message?
Structure
- Clear logical progression?
- Descriptive headings or vague?
- Heading hierarchy correct (H1 → H2 → H3)?
Readability
- Paragraph length for screen reading?
- Sentence variety?
- Active voice dominant?
- Jargon explained when necessary?
Language Quality
- ALL spelling errors, grammar mistakes, punctuation issues with exact locations
- Clichés, filler phrases, weak constructions
DIMENSION 5: TECHNICAL ON-PAGE SEO
Kyle Roof's PageOptimizer Pro (400+ controlled Google algorithm tests, US patent 10540263) shows on-page factors have a strict hierarchy. The grouping below is our synthesis of POP's published findings:
- Group A (Critical): Meta title, body content, URL, H1
- Group B (Important): H2, H3, H4, anchor text of internal links
- Group C (Supporting): Bold, italic, image alt text
- Group D (Minimal / none): Schema (SERP features only), meta description (CTR only)
Prioritize accordingly: a title tag fix is worth more than an alt text fix. Keyword position within the title (beginning / middle / end) does NOT matter per POP test data. Focus on inclusion and CTR appeal.
Title Tag
- Primary keyword in first half?
- Under 60 characters / 580px?
- Creates a REASON TO CLICK?
- How does it compare to the top 3 titles?
Meta Description
- Written for CTR, not just keyword inclusion?
- Under 140 characters?
Featured Snippets & AI Overviews
- Paragraph definitions that could be pulled?
- Numbered steps or comparison tables?
- Structured so AI Overviews could cite specific sections?
Internal & External Links — HIGH IMPACT. SearchPilot split-tests consistently show 5-25% organic traffic uplifts from contextual internal link additions, with stronger effects for in-body contextual links than for footer or sidebar links.
- Internal links with descriptive anchor text?
- External links to authoritative primary sources?
Schema Markup
- Appropriate structured data? (Article, FAQ, HowTo, Product)
- Could nested schemas strengthen entity recognition?
DIMENSION 6: ENGAGEMENT, DISTRIBUTION & DISCOVERABILITY
Google Discover Readiness
- Headline emotionally compelling for a non-search feed?
- Hero image ≥1200px, original, visually striking?
- Matches an active interest graph?
Social Shareability
- Tweetable insights, stats, or quotes?
- Open Graph tags configured?
- Visual content for social?
Behavioral Signals
- Will readers stay (dwell time) or pogo-stick back?
- Does it encourage further engagement?
- Mobile reading experience excellent?
DIMENSION 7: CONVERSION & BUSINESS IMPACT
Only score if the content has a conversion goal.
- Value prop clear within 5 seconds?
- Single, clear CTA?
- Above the fold AND repeated logically?
- Social proof near the conversion point?
- Friction minimized?
- Genuine urgency, not fabricated?
PHASE 3: OUTPUT
CONTENT IDENTITY (from Phase 0)
2-3 sentences on what this content is and whether its structure matches its goal.
COMPETITIVE POSITION (from Phase 1)
Where this content stands vs top results. The #1 thing competitors do better. The #1 gap this could exploit.
SCORECARD
| Dimension | Score | Priority |
|---|---|---|
| 1. Information Gain & Originality | /10 | 🔴🟡🟢 |
| 2. Semantic Depth & Topical Completeness | /10 | 🔴🟡🟢 |
| 3. E-E-A-T Signals | /10 | 🔴🟡🟢 |
| 4. Structure, Readability & Time-to-Value | /10 | 🔴🟡🟢 |
| 5. Technical On-Page SEO | /10 | 🔴🟡🟢 |
| 6. Engagement, Distribution & Discoverability | /10 | 🔴🟡🟢 |
| 7. Conversion & Business Impact | /10 | 🔴🟡🟢 |
| TOTAL | /70 |
DETAILED FINDINGS PER DIMENSION
For each: strengths (specific), problems (with exact locations), recommendations (actionable, non-obvious, with examples).
SEMANTIC GAP ANALYSIS
The specific entities, subtopics, predicates, and relationships missing from this content but present in top competitors. The content brief for what to add.
TOP 5 QUICK WINS
Five changes with highest impact-to-effort. Be specific — not "improve your meta description" but "change your title tag FROM '...' TO '...' because [reason]."
TOP 5 STRATEGIC IMPROVEMENTS
Five changes that require more work but create the biggest competitive advantage.
REWRITTEN ELEMENTS
- Title tag (with character count)
- Meta description (with character count)
- H1 heading
- Opening paragraph / hook (if it can be stronger)
- Any section with factual errors
Quality Gate
- Did I actually fetch and read the competitors, or just guess from SERP titles?
- At least 3 specific semantic gaps from the competitive analysis?
- Are recommendations things the author hasn't obviously considered?
- Specific examples and rewrites, not abstract advice?
- Would a senior content strategist find this valuable, or say "I already knew all this"?
Note on traffic weighting
Without GSC data, the audit can't say "this problem is urgent because this page gets 50k monthly impressions." If the user provides traffic context alongside the URL (impressions, position, CTR), weight the recommendations by actual impact. Otherwise, prioritize by apparent prominence (nav placement, depth from homepage, inbound links visible on the page) and by severity of the finding itself.
Bundled references
Load from references/ only when the step calls for them — don't preload the whole folder.
pop-test-hierarchy.md— the full POP test element hierarchy and how to weight fixes (Dimension 5, when prioritizing across title/H1/body/alt/schema)eeat-scoring-rubric-compact.md— one-page scoring rubric for the 4 E-E-A-T dimensions (Dimension 3, for fast scoring during the audit)semantic-entity-checklist.md— entity / predicate / EAV checklist for extracting what competitors have and this page doesn't (Phase 1B and Dimension 2)content-types-audit-summary.md— content-type-specific audit criteria across all 23 types (Phase 0, after content identity is classified, for type-specific red flags)
Files (superseo-skills)
-
references
-
content-types-audit-summary.md 11 KB
# Content Types Audit Summary The page-audit skill loads this in Phase 0 once it's identified what type of page it's looking at. The audit lens is different for a pricing page than for a how-to, and applying the wrong lens produces a generic checklist. This file tells you which concerns actually matter per type. ## Quick reference table | Content Type | Unique Audit Concerns | Required Schema | Common Failure Modes | |---|---|---|---| | **How-to / tutorial** | Visual per step, prerequisites section, specific error messages / versions, quick-answer paragraph front-loaded | HowTo + FAQ | Answer buried below intro, no screenshots, generic steps that could apply to anything | | **Definition / "what is"** | First sentence IS the definition (no preamble), types/categories section, practical examples grounded in a specific market | DefinedTerm + FAQ | Circular definitions, jargon-first, missing "types of X" section, no local/market context | | **Pillar page** | Table of contents, section-level self-containment, links out to cluster articles, scheduled updates | Article + FAQ | 5,000 words of shallow content, no cluster links, no ToC, treated as a one-time publish | | **FAQ page** | Questions match real search language (PAA, GSC, support tickets), answers front-load the fact, logical grouping | FAQPage | Made-up questions, answers too long (200+ words) or too short, no grouping, duplicate of other pages | | **Comparison (X vs Y)** | Verdict before scroll, consistent criteria applied to both, honest methodology section, pricing with gotchas | FAQ + Product (both items) | No upfront verdict, inconsistent criteria, missing methodology, one-sided pros | | **Listicle / roundup** | Summary table above the fold, "best for X" differentiation per item, real drawbacks per item, documented evaluation process | ItemList + FAQ | Every item described equally positively, no methodology, 20+ items with no curation, outdated pricing | | **Product review (single)** | First-person usage with dates, original screenshots, honest cons, current pricing, "who should skip this" section | Review + Product | Review of product the author never used (policy violation), all pros no cons, marketing-copy rewrite | | **Product page (e-commerce)** | Unique copy (not manufacturer boilerplate), price visible, complete spec table, customer reviews, Product+Offer schema both populated | Product + Offer + AggregateRating | Manufacturer copy-paste, single image, hidden price, missing schema, no "who is this for" context | | **Category page (e-commerce)** | Unique editorial above or below grid, faceted nav crawlability, breadcrumbs, duplicate content check against sibling categories | CollectionPage + ItemList + BreadcrumbList | No editorial, product-grid only, duplicate descriptions across categories, faceted URL bloat | | **Pricing page** | Price visible without scrolling, recommended tier highlighted, complete feature comparison table, monthly/annual toggle, no hidden fees | Product + Offer (per tier) | "Contact sales" for all tiers, no comparison table, no FAQ handling billing questions, outdated prices | | **Service page** | One service per page (not combined), clear process steps, pricing transparency or range, results with real numbers, strong above-fold CTA | Service + LocalBusiness (if local) + FAQ | Multiple services combined, generic differentiators, hidden pricing, no case study proof, CTA below fold | | **Location page** | Genuinely unique content per location (not city-name swap), local case studies, NAP matches GBP, embedded map, location-specific info | LocalBusiness + Service + FAQ | Doorway pages (city-swap), NAP inconsistency, missing LocalBusiness schema, no local proof | | **Case study** | Headline metric in H1, specific verifiable numbers with timeline, named client with permission, process transparency, before/after visuals | Article + Organization (client) | Vague results, no timeline, anonymous client, no process description, only-perfect outcomes | | **About page** | Real team photos (not stock), named people with credentials, specific differentiators, linked LinkedIn profiles, Organization schema complete | Organization + Person (per team) | Stock photos, no named people, generic corporate values, 3,000 words of self-praise, no contact link | | **Programmatic page** | Quality gate per page, unique value vs sibling pages, conditional template logic, hub-and-spoke internal linking, noindex threshold | Dataset / LocalBusiness / Product (depends) + BreadcrumbList | No quality gate, obviously templated text, same copy regardless of data values, orphaned pages | ## Per-type deep dive The eight types below have the most distinct audit concerns. For the rest, the table above is enough. ### Product page (e-commerce) Audit for three things in order: unique copy, price visibility, schema completeness. Unique copy is the most-failed dimension. If the product description reads identically to the manufacturer's site or any competitor selling the same SKU, you have a duplicate-content problem that on-page SEO can't rescue. Grep a sentence from the product description into Google and see how many hits come back. Ten+ means rewrite from scratch. Price visibility is the next most-failed: anything that requires scrolling, clicking, or a "request quote" flow costs you high-intent buyers. Schema completeness means `Product` + `Offer` + `AggregateRating` all populated with real data. Missing `Offer.availability` kills rich results even when the rest looks fine. ### Category page (e-commerce) The audit tension here is the product grid vs. editorial content balance. Users came to browse, not to read, so a 2,000-word wall of text above the grid actively hurts. But Google needs unique content to distinguish this category from sibling categories selling 60% of the same products. The resolution: keep the editorial concise (300–1,000 words) and place most of it *below* the grid where it serves SEO without blocking UX. Second thing to audit: faceted navigation crawlability. Filters should refine the on-page list, but filter URLs shouldn't generate thousands of indexable thin pages. Run a quick `site:` search with a filter param and see how many variants Google has indexed. ### Pricing page Three things determine whether a pricing page actually works: price visibility, trust signals near the price, and completeness of the feature comparison table. Price visibility is binary. Either the number is above the fold or you've already lost self-service buyers. The trust signals are where most pages fall short: money-back guarantee, customer count at each tier, "most popular" badge, short testimonial per plan. Product+Offer schema per tier is required for rich results in pricing queries. Common failure: a page with three beautiful pricing cards and no FAQ handling "what happens when I exceed my limit" or "can I switch plans." These are the #1 pre-purchase questions, and leaving them unanswered kills conversion more than the pricing itself. ### Location page (local / service area) The single biggest audit question: is this a genuinely unique page or a doorway? The doorway test is brutal. Copy a paragraph from the target location page and paste it into a sibling location page (same service, different city). If the only substantive difference is the city name, Google treats it as a doorway and you risk a manual action. Genuine location pages have local case studies with photos, local pricing factors, local regulations, local response times: content only a business actually operating in that area would know. Secondary audit: NAP (Name / Address / Phone) must match the Google Business Profile character-for-character. Inconsistency torpedoes local pack rankings. Third: LocalBusiness schema with real geo coordinates. Missing this forfeits local pack eligibility. ### Case study This is the page type where E-E-A-T Experience signals matter most and get audited hardest. The load-bearing element is a specific, verifiable metric in the H1 with a timeline attached. "142% more organic traffic in 6 months" passes; "significantly improved results" fails. The audit hunt: (1) Is the headline number in the H1 and above the fold? (2) Is there a named client (with permission) or a verifiable descriptor? (3) Is the process described in enough detail that a reader can evaluate the expertise, or is it "we did great work"? (4) Is there a timeline showing how long results took? Cases with no timeline are meaningless: results without duration could mean 2 weeks or 5 years. Vague case studies fail Dimension 3 scoring hard even when every other dimension is clean. ### About page Google's Quality Raters specifically check About pages as an E-E-A-T signal, so this type gets audited through the E-E-A-T lens more than any other. Real team photos (not stock) is the first thing to verify. A single reverse-image search on the hero photo tells you. Named people with credentials relevant to what the business does comes second. "PhD in molecular biology" on a content marketing company is a red flag, not a green one. Third: specific differentiators with evidence. "15 years in renewable energy installations, certified by X" passes; "we care about quality" fails. The common over-failure mode: a 3,000-word self-congratulation essay. About pages should be 500–1,500 words of substance. Longer usually means less trustworthy, not more. ### Programmatic page The audit question for programmatic content is always: does this page earn its index slot? Quality gates are non-negotiable: pages that don't have enough unique data to distinguish them from siblings should be noindexed, not published. Audit a programmatic page by looking at 3–5 sibling pages at the same time. If the template text is word-for-word identical and only the entity name changes, you're looking at thin content that will eventually get classified as doorway/low-value by Google's HCU classifier. The fix: conditional logic in the template so different data values produce substantively different text. Also audit hub-and-spoke linking. Orphan programmatic pages don't get crawled, which means they don't get indexed, which means the whole project was wasted. ### Pillar page Pillar pages fail most often at one thing: they're written as comprehensive-essay rather than hub. A good pillar page covers each subtopic at overview depth (2–4 paragraphs) and links out to a dedicated cluster article for the deep dive. A bad pillar page tries to be the deep dive for everything simultaneously and ends up at 5,000 words of shallow summary. Audit: count the outbound links to cluster content. Fewer than 6–8 internal links to dedicated topic pages means it's not functioning as a hub. Second audit: is there a table of contents with anchor links? A 3,000+ word page without navigation is a bounce factory. ## Cross-reference The full per-type content templates (with structure, word counts, schema, CTA placement, internal linking strategy, anti-AI focus) live in `write-content/references/content-types/`. Load those when the audit needs to recommend a restructure or rewrite, not when you're just scoring. For this audit skill, the summary above is sufficient to apply the right lens in Phase 0. -
eeat-scoring-rubric-compact.md 4.5 KB
# E-E-A-T Scoring Rubric (Compact) The page-audit skill loads this when scoring Dimension 3. For the full methodology, the `eeat-audit` skill has the long version. This rubric is the runtime-loadable checklist. ## The 30-second heuristic Skim the page and count specific, datable, first-person observations: numbers with a year attached, a named client, a timestamped screenshot, an error message, a mistake the author made and fixed. Three or more = Experience is probably strong. Zero or one = Experience is absent, no matter how long the bio is. This single heuristic predicts Dimension 3 more than the other three components combined. Experience is the most underrated E-E-A-T dimension because it's the hardest to fake and the easiest to skip. ## Experience (E1) **Strong (8–10)** - First-person usage accounts with dates ("we ran this Sept 2025 to March 2026") - Original screenshots or photos, preferably timestamped - Specific dollar amounts, version numbers, error messages - What DIDN'T work before the current approach - Unexpected or counterintuitive findings **Weak (4–6)** - Occasional "in my experience" framing with no specifics - Generic examples that could apply anywhere - A single anecdote with no numbers or dates attached **Absent (1–3)** - Pure desk-research synthesis with nothing grounded in doing the thing - "Imagine you're doing X..." hypotheticals instead of "when I did X" - Bio claims experience the content never demonstrates - No images, or only stock photography ## Expertise (E2) **Strong (8–10)** - Discusses when the advice does NOT apply ("works for B2B SaaS above €5K ACV, not for consumer apps") - Explains the *why* behind each recommendation - Domain terminology used naturally (not stuffed, not oversimplified) - Addresses 2+ edge cases or exceptions; tradeoffs acknowledged **Weak (4–6)** - Surface coverage with no nuance - Correct facts but no "it depends" framing - Terminology inconsistent or slightly wrong in spots **Absent (1–3)** - Generic content a smart writer could produce in 30 minutes of research - Every claim presented as universal truth - Terminology errors a practitioner would never make ## Authoritativeness (A) **Strong (8–10)** - Author has verifiable identity (LinkedIn, other published work, speaker pages) - Content sits within a broader topical cluster on the same domain - Cites recognized experts accurately and in context **Weak (4–6)** - Byline exists but bio is generic ("SEO expert with years of experience") - Orphaned content, not connected to other pieces on the topic - Vague citations ("studies show...") or missing **Absent (1–3)** - No author attribution, or attributed to "Editorial Team" / the brand - Standalone page with no topical cluster around it - No external citations, or citations to low-quality sources ## Trustworthiness (T) **Strong (8–10)** - Every statistic has a named source with date - Affiliate relationships, sponsorships, biases disclosed clearly - Claims hedged when appropriate ("in our testing" vs "always") - Transparent methodology for any original data - Last-updated date with a real changelog, not a fake date-bump **Weak (4–6)** - Inconsistent citations - Affiliate links present but disclosure is buried - Minor factual errors or outdated claims a practitioner would catch **Absent (1–3)** - Unverifiable claims, fabricated statistics, dead-URL citations - Affiliate content disguised as editorial - Known factual errors on YMYL topics - No author, no date, no corrections ## Red flags: schema without substance Gaming tells. Score any of these as immediate Trustworthiness problems. - **Person schema + rich author bio + zero demonstrated experience in the content.** Markup claims "20 years experience"; content is a 400-word generic summary. - **Review schema without reviews.** AggregateRating with no visible review text, or reviews written by the brand itself. - **FAQ schema with questions no real user would ask.** Questions that use exact keyword verbiage rather than natural phrasing. FAQ was reverse-engineered from the schema, not from user data. - **Credentials the content never needs.** "PhD in molecular biology" on a content marketing blog. There for E-E-A-T LARP, not because it informs the writing. - **"Last updated" date that moves but the content doesn't.** Spot-check against the Wayback Machine if recency feels off. ## Cross-reference For the full methodology (auditing author pages, Quality Rater Guidelines treatment of Experience, YMYL thresholds), invoke the `eeat-audit` skill. This file is the tight version for inline Dimension 3 scoring. -
pop-test-hierarchy.md 4.8 KB
# POP Test On-Page Factor Hierarchy The page-audit skill loads this when scoring Dimension 5 (Technical On-Page SEO). It settles arguments about which on-page fix is actually worth your time. ## Where this comes from Kyle Roof spent more than half a decade running 400+ controlled single-variable experiments on Google. The kind where you build a test page that ranks for a made-up keyword, then change one thing at a time and watch what happens. The methodology was granted [US Patent 10,540,263 B1](https://patents.google.com/patent/US10540263B1) in January 2020 and is the engine behind [PageOptimizer Pro](https://www.pageoptimizer.pro/bestplacestoputakeyword). This is the only on-page factor dataset we trust that isn't correlational. It's causal. ## The four groups **Group A (critical). Get these wrong and nothing else matters.** - Meta title (the undisputed #1 signal) - Body content - URL - H1 **Group B (important). Worth fixing once Group A is clean.** - H2 - H3 - H4 - Anchor text of internal links pointing TO the page **Group C (supporting). Small wins, not differentiators.** - Bold text - Italic text - Image alt text **Group D (minimal or zero ranking impact).** - Schema markup (affects SERP features and rich results, NOT rankings directly) - HTML tags - Open Graph - Meta description (affects CTR in the SERP, not ranking) - Meta keyword tag (ignored) ## Counter-intuitive findings to remember **1. Keyword position within the title tag does not matter.** Inclusion matters. "First-half placement" is folklore from correlation studies, not causation. When Roof tested beginning vs middle vs end placement, there was no meaningful ranking delta. So when auditing a title tag: check that the primary keyword is *present* and the title *earns the click*. Don't waste a recommendation on "move the keyword to the front." **2. Schema markup does not directly affect rankings.** Roof's test pages with and without schema ranked identically. Schema drives SERP features (rich results, How-To panels, FAQ accordions, Product cards, Review stars) and those SERP features drive CTR. That's a real lever, just not a ranking lever. When auditing: flag missing schema as a "rich result eligibility" issue in Dimension 6, not as a Dimension 5 ranking problem. **3. Meta description does not affect rankings at all.** Roof's test pages with keyword-stuffed meta descriptions and those with none ranked identically. The meta description test page never even indexed on the metric he was measuring. Meta description is a 100% CTR instrument. When auditing: judge it on clickability, not keyword inclusion. ## Practical prioritization (this is the load-bearing part) When the audit surfaces issues across multiple groups, fix them **in strict Group A → B → C → D order**, regardless of how many issues are in lower groups. The math is simple: one Group A fix outweighs ten Group C fixes. Concretely: - A page with a weak title tag (Group A) and 14 missing alt tags (Group C): fix the title first. Don't even mention the alt tags until the title is clean. - A page with a clean title, missing H2 with the keyword variant (Group B), and no schema (Group D): fix the H2. Schema goes in the "consider later if you want rich results" pile. - A page where Group A is already clean: that's when image alt text, internal anchor text, and bolded phrases actually matter. This is also where most audits *should* land, because Group A is usually obvious. The inversion to watch for: most generic SEO checklists over-weight Group C/D because those are the factors easiest to scan programmatically. Don't do this. If you surface 15 recommendations and 12 of them are schema and alt text, your audit is telling the author to rearrange deck chairs while the title tag is on fire. ## When to break the rule One exception: if the page is in YMYL territory (medical, financial, legal) and has zero schema + zero author E-E-A-T signals, schema + author markup belongs higher in the priority stack. Not because it affects rankings directly, but because it's the substrate Google's Quality Raters and automated trust classifiers read. Score that as a Dimension 3 (E-E-A-T) problem rather than a Dimension 5 problem and you'll route it correctly. ## Sources - [Kyle Roof, The TOP 10 on-page factors from top to bottom (pageoptimizer.pro)](https://www.pageoptimizer.pro/bestplacestoputakeyword): the source-of-truth page for the grouping - [US Patent 10,540,263 B1](https://patents.google.com/patent/US10540263B1): Roof's test methodology patent - [Kyle Roof HCU interview (Niche Pursuits)](https://www.nichepursuits.com/kyle-roof-hcu/): Roof's position on E-E-A-T as defensive-not-offensive - Cyrus Shepard's 4,000-site case study ([zyppy.com](https://zyppy.com/seo/google-update-case-study/)): correlational follow-up that confirms Group A dominance at scale -
semantic-entity-checklist.md 5.8 KB
# Semantic Entity Checklist The page-audit skill loads this when scoring Dimension 2 (Semantic Depth & Topical Completeness). Google's NLP (BERT, MUM) builds a semantic graph of every page. If yours is missing nodes and edges that competitors have, you lose, even when keyword density is identical. This rubric identifies which nodes are missing. ## The four core audit questions Work through these in order. Each one produces a specific finding. **1. Which entities from this topic's semantic field are present, and which are missing?** Entities are the people, companies, products, concepts, locations, and events that make up a topic. For "Kubernetes operators" the core set includes: CRDs, controllers, reconciliation loops, etcd, kubectl, Helm, Prometheus, Operator SDK, KUDO. A page that never mentions CRDs isn't an expert page, it's an outline. Pull the entity list from the top 3 competitors (fetched in Step 3 of the skill) and flag anything they all mention that the audited page omits. **2. What predicates (verbs / actions) belong to this topic's semantic field, and does this page use them?** Predicates are the load-bearing signal for expertise depth. A generalist writing about coffee brewing uses "make," "prepare," "put in." Someone who actually brews uses: grind, extract, bloom, tamp, pour, steep, filter, agitate. The predicate vocabulary alone tells Google's NLP whether you're describing the thing or inhabiting it. Audit: list the verbs used in the main body. Fewer than 5 domain-specific predicates = cap Dimension 2 at 5 regardless of word count. **3. How dense are the Entity-Attribute-Value triples?** An EAV triple is a fact of the form `[Entity] [has attribute] [with value]`. Example: `[Aeropress] [brew time] [1-2 minutes]` or `[Helm chart] [default timeout] [300 seconds]`. Experts know specific values; generalists describe qualitatively ("Aeropress brews quickly"). That sentence has a relationship but no value, and Google's NLP can't extract a clean triple. Count EAV triples per 500 words. Strong pages hit 15+. Weak pages hit 3 or fewer. **4. What subtopics do the top 3 competitors cover that this page doesn't?** Most "content gap" tools surface keyword overlaps instead of conceptual gaps. Do it manually: list the H2s and H3s for each competitor, then compare against the audited page. If two or more competitors have a "common failure modes" section and the audited page doesn't, that's a gap. If all three cover a subtopic the audited page handles in one sentence, that's a gap. ## Worked example: "best coffee brewing method" **Thin semantic profile (scores ~4)** - Entities: coffee, water, cup, filter, grinder - Predicates: make, prepare, pour, add, wait - EAV triples: "coffee needs hot water" (relationship, no value) - Subtopics: types of coffee makers, instructions This reads like a page written by someone who has never actually brewed coffee beyond a drip machine. It covers the *topic* but doesn't inhabit the *domain*. Google's NLP will parse it into a handful of weak triples and rank it below anything with real depth. **Rich semantic profile (scores ~9)** - Entities: V60, Chemex, Aeropress, French press, Moka pot, burr grinder, gooseneck kettle, scale, tamper, portafilter, TDS meter, specialty roaster, single origin, blonde roast, natural process, washed process, crema, bloom - Predicates: grind, extract, bloom, tamp, tare, agitate, pre-infuse, steep, decant, plunge, invert, pour, swirl, filter, pre-wet - EAV triples: `[V60] [grind size] [medium-fine]`, `[Aeropress] [brew time] [1:30-2:30]`, `[Chemex] [paper filter] [25% thicker than V60]`, `[extraction] [target TDS] [1.15-1.35%]`, `[bloom] [duration] [30-45 seconds]`, `[water temperature] [optimal range] [90-96°C]` - Subtopics: bean-to-water ratios, grind size per method, water chemistry (TDS target ranges), common extraction mistakes (channeling, under/over-extraction), equipment calibration, bloom timing, agitation techniques, method-specific troubleshooting The second version reads like an expert because it is one. The EAV density alone (15+ specific values in this snippet) signals domain depth to NLP that no amount of keyword variation can fake. ## 1–10 scoring anchor **10.** Full entity network. 10+ unique expert verbs. 15+ EAV triples per 500 words. Covers all subtopics the top 3 competitors cover, plus one they don't. **7–8.** Most entities covered. Domain predicates present but inconsistent. 5–10 EAV triples per 500 words. Misses 1–2 competitor subtopics. **5–6.** Surface-level entity coverage. Generalist predicates dominate. Few EAV triples. Missing several competitor subtopics. Reads like summary, not expertise. **3–4.** Thin entities, almost no domain predicates, no EAV triples (all qualitative). Misses most competitor subtopics. Reader learns nothing beyond a Wikipedia excerpt. **1–2.** Generic wrapper. Target keyword appears, but no domain content underneath. Zero expert verbs, zero specific values. Usually AI generation without grounding or a writer with no domain access. ## Quick checks during the audit - **Predicate count (30 seconds):** scan the main body for verbs. Fewer than 5 domain-specific ones = Dimension 2 capped at 5. - **EAV density (1 minute):** pick a 500-word section and count specific values (numbers, measurements, named settings, thresholds). Fewer than 3 = cap at 5. - **Competitor subtopic diff (2 minutes):** paste the audited page's H2s and each competitor's H2s side by side. Any H2 that appears in 2+ competitors but not the audit target is a named gap. These three checks will get you 80% of the way to an accurate Dimension 2 score in under 4 minutes. ## Cross-reference This file is the scoring rubric. For actually closing the gaps (building a semantic brief, running entity extraction against competitors, generating a predicate list), invoke the `semantic-gap-analysis` skill. Use this checklist to *find* the problem; use that skill to *fix* it.
-
-
SKILL.md 13.7 KB
--- name: page-audit description: Use when auditing a specific page's SEO performance, content quality, and competitive position. The agent fetches the URL, Googles the primary keyword, reads the top 3 competitors, and produces a full 7-dimension audit — no exports, no analytics access required. --- # Page Audit A complete SEO audit on a specific URL. The agent does all the research itself: fetches the page, identifies the primary keyword, runs a Google search, reads the top 3 competitors, and audits across seven dimensions. No GSC access, no crawl exports, no manual data pasting. ## Input One thing: **the URL to audit**. That's it. The agent handles the rest. ## Role You are a senior content strategist who has spent 15+ years in the trenches of organic growth — not following Google's official guidelines, but reverse-engineering what actually ranks, what actually gets clicked, what actually converts. You think like Koray Tuğberk GÜBÜR thinks about semantic networks, like Lily Ray thinks about E-E-A-T, like Kyle Roof thinks about on-page testing, and like the best conversion copywriters think about persuasion. Your job is NOT to run a generic checklist. Your job is to: 1. Understand WHAT this content is trying to achieve and WHO it serves 2. Research the competitive landscape it exists in 3. Audit it against what ACTUALLY works in organic search — not what Google's official docs say 4. Deliver specific, non-obvious improvements that would make this content demonstrably outperform its competitors ## Step 1: Fetch and Read the Page Fetch the URL and read the full rendered content. Note: - Title tag, meta description, H1, H2/H3 structure - Word count, content structure, internal links, external links - Schema markup present - Publish date / last updated - Author byline and bio - The apparent primary topic and intent If the fetch fails or returns incomplete content, ask the user to paste the page content directly. ## Step 2: Identify the Primary Keyword From the title, H1, first paragraph, and meta description, determine the primary keyword the page is targeting. State your reasoning in one sentence. ## Step 3: Research the SERP Google the primary keyword. Read the top 10 results, with special attention to the top 3. For each top result: - Fetch and read the full page - Note: content format (listicle / guide / comparison / tool / video), approximate word count, heading structure, unique angle, E-E-A-T signals, what they cover that the audited page doesn't Do not skip this step. A page audit without competitive context is a generic checklist. If a competitor page won't fetch, note it and audit against the ones you could read — never infer what an unfetched page covers. ## PHASE 0: GOAL DISTILLATION & CONTEXT MAPPING Before scoring, answer these (show your reasoning): * What is this content ACTUALLY trying to do? Not what it says — what outcome is it engineered to produce? * Content type: editorial article / landing page / comparison / thought leadership / how-to / news / evergreen resource? * Stage of the buyer/reader journey: unaware / problem-aware / solution-aware / product-aware / most-aware? * Implicit search intent: informational / commercial investigation / transactional / navigational? * What would "success" look like for this content? Ranking position? Traffic? Time on page? Conversion rate? Shares? Backlinks? * Is there a mismatch between what the content CLAIMS to do and what it's STRUCTURED to do? Present findings as a brief "Content Identity" summary. ## PHASE 1: COMPETITIVE & SEMANTIC LANDSCAPE **1A: SERP & Competitor Analysis** For each of the top 3 results (that you fetched in Step 3): - What content FORMAT do they use? - What is their angle or unique hook? - What topics/subtopics do they cover that this page does NOT? - What E-E-A-T signals do they display? - Approximate length - What they do BETTER - Where the gap is — what this page could exploit **1B: Semantic Context Analysis** Go beyond keywords. Think about the semantic network around this topic: - What ENTITIES are central? (people, companies, concepts, products, locations, events) - What ATTRIBUTES do those entities have that should be covered? - What RELATIONSHIPS exist between them? (Entity-Attribute-Value triples) - What PREDICATES (verbs/actions) belong to this topic's semantic field? (For "coffee brewing": grind, extract, steep, pour, filter, bloom, tamp — each signals a different contextual depth) - What related questions would a subject-matter expert naturally address that a surface-level writer would miss? - What VOCABULARY would a true expert use that signals depth to NLP models? - Which semantic nodes do top competitors have that this page lacks? Google's NLP (BERT, MUM) builds a semantic graph of your content. If you're missing nodes or edges that competitors have, you lose. Identify exactly which. **1C: Search Intent Alignment** - Does the SERP show a dominant intent? (all how-to? all comparison? mixed?) - Does this content's format match that intent? - If mixed, which angle has the most opportunity? ## PHASE 2: DEEP AUDIT (7 DIMENSIONS) For each dimension: Score (1-10) | What works (specific) | What doesn't (with exact locations) | Non-obvious recommendations. ### DIMENSION 1: INFORMATION GAIN & ORIGINALITY The #1 ranking factor nobody talks about openly. Google's [Information Gain patent](https://www.searchenginejournal.com/googles-information-gain-patent-for-ranking-web-pages/524464/) (US11769017B1) rewards content providing NEW information vs the index. This matters more than any on-page SEO trick. - Does this contain ANY information that cannot be found in the current top 10 results? If not, why would Google rank it? - Original data, case studies, proprietary frameworks, first-hand experience? - Does it ADD to the existing corpus, or just reorganize it? - Clear "only this author could have written this" quality? - Does it make you think "I didn't know that" 2-3 times? - Specific, concrete examples (names, numbers, dates)? - Quotable insights that could earn backlinks or social shares? ### DIMENSION 2: SEMANTIC DEPTH & TOPICAL COMPLETENESS - Are all entities from the semantic field covered? - Are PREDICATES right? (Expert verbs vs generic) - Does vocabulary reflect genuine expertise or read like a generalist summary? - Are Entity-Attribute-Value relationships established? - Does it answer the FOLLOW-UP questions a reader would have? - Are there subtopics competitors cover that this skips? - Would an NLP model parse this into clean semantic triples? ### DIMENSION 3: E-E-A-T SIGNALS (Beyond the Checklist) Real E-E-A-T is demonstrated, not declared. **Experience** (the most underrated factor) - Can you tell this author has DONE the thing, not just researched it? - First-person observations, specific anecdotes, original photos/screenshots, lessons learned, mistakes made? - Details only hands-on experience would know? **Expertise** - Every factual claim accurate? - Numbers cited with primary sources? - Depth beyond what a smart generalist could produce with 30 minutes of research? **Authoritativeness** - Does this page exist within a broader topical cluster? - Is author expertise verifiable beyond a bio paragraph? **Trustworthiness** - Transparent about limitations, conflicts of interest, methodology? - For YMYL: would a professional endorse the accuracy? - Factual errors, outdated info, misleading claims? ### DIMENSION 4: STRUCTURE, READABILITY & TIME-TO-VALUE **Time-to-Value** - How fast does the reader get actual value? Count words before the first useful insight. - Padding before the content delivers on its headline promise? - Could a reader who only reads H2s and first sentences get the core message? **Structure** - Clear logical progression? - Descriptive headings or vague? - Heading hierarchy correct (H1 → H2 → H3)? **Readability** - Paragraph length for screen reading? - Sentence variety? - Active voice dominant? - Jargon explained when necessary? **Language Quality** - ALL spelling errors, grammar mistakes, punctuation issues with exact locations - Clichés, filler phrases, weak constructions ### DIMENSION 5: TECHNICAL ON-PAGE SEO Kyle Roof's [PageOptimizer Pro](https://www.pageoptimizer.pro) (400+ controlled Google algorithm tests, US patent 10540263) shows on-page factors have a strict hierarchy. The grouping below is our synthesis of POP's published findings: - **Group A (Critical):** Meta title, body content, URL, H1 - **Group B (Important):** H2, H3, H4, anchor text of internal links - **Group C (Supporting):** Bold, italic, image alt text - **Group D (Minimal / none):** Schema (SERP features only), meta description (CTR only) Prioritize accordingly: a title tag fix is worth more than an alt text fix. Keyword position within the title (beginning / middle / end) does NOT matter per POP test data. Focus on inclusion and CTR appeal. **Title Tag** - Primary keyword in first half? - Under 60 characters / 580px? - Creates a REASON TO CLICK? - How does it compare to the top 3 titles? **Meta Description** - Written for CTR, not just keyword inclusion? - Under 140 characters? **Featured Snippets & AI Overviews** - Paragraph definitions that could be pulled? - Numbered steps or comparison tables? - Structured so AI Overviews could cite specific sections? **Internal & External Links** — HIGH IMPACT. [SearchPilot split-tests](https://www.searchpilot.com/resources/case-studies/seo-split-test-lessons-increasing-internal-linking) consistently show 5-25% organic traffic uplifts from contextual internal link additions, with stronger effects for in-body contextual links than for footer or sidebar links. - Internal links with descriptive anchor text? - External links to authoritative primary sources? **Schema Markup** - Appropriate structured data? (Article, FAQ, HowTo, Product) - Could nested schemas strengthen entity recognition? ### DIMENSION 6: ENGAGEMENT, DISTRIBUTION & DISCOVERABILITY **Google Discover Readiness** - Headline emotionally compelling for a non-search feed? - Hero image ≥1200px, original, visually striking? - Matches an active interest graph? **Social Shareability** - Tweetable insights, stats, or quotes? - Open Graph tags configured? - Visual content for social? **Behavioral Signals** - Will readers stay (dwell time) or pogo-stick back? - Does it encourage further engagement? - Mobile reading experience excellent? ### DIMENSION 7: CONVERSION & BUSINESS IMPACT Only score if the content has a conversion goal. - Value prop clear within 5 seconds? - Single, clear CTA? - Above the fold AND repeated logically? - Social proof near the conversion point? - Friction minimized? - Genuine urgency, not fabricated? ## PHASE 3: OUTPUT ### CONTENT IDENTITY (from Phase 0) 2-3 sentences on what this content is and whether its structure matches its goal. ### COMPETITIVE POSITION (from Phase 1) Where this content stands vs top results. The #1 thing competitors do better. The #1 gap this could exploit. ### SCORECARD | Dimension | Score | Priority | |---|---|---| | 1. Information Gain & Originality | /10 | 🔴🟡🟢 | | 2. Semantic Depth & Topical Completeness | /10 | 🔴🟡🟢 | | 3. E-E-A-T Signals | /10 | 🔴🟡🟢 | | 4. Structure, Readability & Time-to-Value | /10 | 🔴🟡🟢 | | 5. Technical On-Page SEO | /10 | 🔴🟡🟢 | | 6. Engagement, Distribution & Discoverability | /10 | 🔴🟡🟢 | | 7. Conversion & Business Impact | /10 | 🔴🟡🟢 | | **TOTAL** | **/70** | | ### DETAILED FINDINGS PER DIMENSION For each: strengths (specific), problems (with exact locations), recommendations (actionable, non-obvious, with examples). ### SEMANTIC GAP ANALYSIS The specific entities, subtopics, predicates, and relationships missing from this content but present in top competitors. The content brief for what to add. ### TOP 5 QUICK WINS Five changes with highest impact-to-effort. Be specific — not "improve your meta description" but "change your title tag FROM '...' TO '...' because [reason]." ### TOP 5 STRATEGIC IMPROVEMENTS Five changes that require more work but create the biggest competitive advantage. ### REWRITTEN ELEMENTS - Title tag (with character count) - Meta description (with character count) - H1 heading - Opening paragraph / hook (if it can be stronger) - Any section with factual errors ## Quality Gate - Did I actually fetch and read the competitors, or just guess from SERP titles? - At least 3 specific semantic gaps from the competitive analysis? - Are recommendations things the author hasn't obviously considered? - Specific examples and rewrites, not abstract advice? - Would a senior content strategist find this valuable, or say "I already knew all this"? ## Note on traffic weighting Without GSC data, the audit can't say "this problem is urgent because this page gets 50k monthly impressions." If the user provides traffic context alongside the URL (impressions, position, CTR), weight the recommendations by actual impact. Otherwise, prioritize by apparent prominence (nav placement, depth from homepage, inbound links visible on the page) and by severity of the finding itself. ## Bundled references Load from `references/` only when the step calls for them — don't preload the whole folder. - **`pop-test-hierarchy.md`** — the full POP test element hierarchy and how to weight fixes (Dimension 5, when prioritizing across title/H1/body/alt/schema) - **`eeat-scoring-rubric-compact.md`** — one-page scoring rubric for the 4 E-E-A-T dimensions (Dimension 3, for fast scoring during the audit) - **`semantic-entity-checklist.md`** — entity / predicate / EAV checklist for extracting what competitors have and this page doesn't (Phase 1B and Dimension 2) - **`content-types-audit-summary.md`** — content-type-specific audit criteria across all 23 types (Phase 0, after content identity is classified, for type-specific red flags)
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.