Research · Dry run

A small test of query wording and search results.

On 19 July we ran ten Indonesian queries through a new SERP collector. This page is the test log: what it captured, where it failed, and what we would change before using it for research.

10Indonesian queries
12pairwise comparisons
1session — no retest baseline
9/12comparisons incomplete

For SEO writers

The short version

People can search for the same topic while asking for different things. “What is NPWP?” needs a definition. “How do I get an NPWP?” needs steps. “Who must have an NPWP?” needs eligibility rules.

In this small test, queries asking for the same kind of answer tended to share more search results than queries asking for different kinds of answers. The test was too small and too messy to prove a ranking effect.

Writing takeaway: identify the answer format behind the query, then make that answer easy to find on the page. Treat this as a useful editorial habit, not a ranking hack.

Why we ran it

We needed to know whether the collector could normalise and compare Google Indonesia result sets. Ten queries were enough to test that machinery, but not enough to study rankings.

A 50-query pilot was next on the list. We held it back after this run exposed three gaps: no repeat captures, no record of SERP module boundaries, and no retry when a result block ended early.

The raw table is public because those gaps are useful debugging information.

The query pairs

Indonesian question words often signal the kind of answer someone wants:

We wanted to see whether the collector noticed when a wording change came with a different result set. The awkward part is that wording and intent often change together. A pair that moves from “how” to “why” does more than swap grammar; it asks for another kind of answer.

Method

Ten Indonesian queries across three topics: removing ads from a phone, BPJS membership status, and NPWP. Within each topic the wording was varied to produce grammatical or intent contrasts.

We compared each query pair in three ways: shared URLs, meaning results found in both searches; Jaccard similarity, the share of unique results common to both searches; and mean rank shift, the average position change for results that appeared in both. A higher Jaccard number means the two result sets looked more alike.

What went wrong

  1. No test-retest baseline. No query was captured twice, so no overlap value on this page can be separated from ordinary session-to-session SERP churn. We do not know what Jaccard the same query scores against itself. Without that number, "high overlap" and "low overlap" have no reference point.
  2. No unrelated-query control. We never measured the overlap between two arbitrary, unrelated queries. A Jaccard of 0.000 shows disjoint sets but gives no measure of significance.
  3. Nine of twelve pairs are incomplete. Five of the ten queries returned fewer than ten results because the parser reached the end of the visible block. Those rows are flagged * throughout and should be read as overlap of the observed sets, not a top-10 comparison.
  4. Unequal set sizes distort Jaccard. A pair where one query returned 7 results has a smaller union, which inflates the ratio relative to a 10-versus-10 pair. Raw Jaccard values are not comparable across rows. See the worked example below.
  5. SERP modules were not separated. Video, social, organic, and other result types were pooled.
  6. Several treatments changed grammar and intent together. Those pairs cannot isolate a grammar effect even in principle.
  7. No page content was analysed. The pipeline collected URLs. It never fetched a ranking page, so this run says nothing about what those pages contained.
  8. No controls for anything else. No matched pages were evaluated for backlinks, authority, freshness, speed, or any other ranking variable.

We ended up with one capture, missing results, and no baseline for ordinary SERP churn.

The ten queries

IDQueryAnswer type requestedResults observedAI Overview in this capture
A1cara menghilangkan iklan di hpProcedure10Yes
A2bagaimana cara menghilangkan iklan di hpProcedure10Yes
A3cara menghilangkan iklan di hp android yang tiba-tiba muncul tanpa aplikasiConstrained procedure10Yes
B1cara cek bpjs aktif atau tidakVerification procedure8 *Yes
B2cara mengecek bpjs aktif atau tidakVerification procedure10Yes
B3kenapa bpjs tidak aktifCausal explanation8 *Yes
N1apa itu npwpDefinition7 *Yes
N2Bagaimana cara punya NPWP?Procedure10Yes
N3apa saja persyaratan membuat npwpRequirements9 *Yes
N4siapa yang wajib memiliki npwpEligibility / obligation9 *Yes

* Fewer than ten results: the parser reached the end of the visible result block. AI Overview appeared for all ten queries in this single capture. That is a snapshot, not a prevalence estimate.

All twelve pairwise comparisons

The table includes every pair produced by the run.

PairContrastSet sizes Shared top 5Jaccard (top 5) Shared (observed sets)Jaccard (observed sets) Mean rank shift
A1–A2carabagaimana cara10 v 1030.42970.5381.00
A1–A3baseline → constrained procedure10 v 1020.25030.1762.33
A2–A3interrogative → constrained procedure10 v 1030.42950.3331.60
B1–B2cekmengecek8 v 10 *40.66770.636 *1.14
B1–B3procedure → cause8 v 8 *00.00000.000 *
B2–B3procedure → cause10 v 8 *00.00000.000 *
N1–N2definition → acquisition7 v 10 *10.11110.063 *1.00
N1–N3definition → requirements7 v 9 *00.00010.067 *5.00
N1–N4definition → obligation7 v 9 *00.00010.067 *4.00
N2–N3acquisition → requirements10 v 9 *00.00020.118 *7.50
N2–N4acquisition → obligation10 v 9 *00.00010.056 *3.00
N3–N4requirements → obligation9 v 9 *00.00020.125 *5.00

* At least one member returned fewer than ten observed results. Read as overlap of the observed sets, not a complete top-10 comparison. Mean rank shift is undefined where no URLs are shared.

The same twelve pairs, grouped by what the wording change actually did

Each dot is one pair. Hollow dots mark comparisons where at least one query returned fewer than ten results. Intent-preserving pairs sit farther right; pairs that change the requested answer type cluster near zero. With no test-retest baseline, the axis has no noise reference.

Jaccard overlap of twelve query pairs, grouped by contrast type Two pairs that changed grammar while holding intent constant scored 0.538 and 0.636. Two pairs that added specificity scored 0.176 and 0.333. Eight pairs that changed the requested answer type scored between 0.000 and 0.125. No test-retest baseline was measured, so the noise floor is unknown. The full figures are in the table above this chart. No test–retest baseline measured — the noise floor could sit anywhere here Grammar changed, intent held · 2 pairs Specificity added · 2 pairs Answer type changed · 8 pairs 0.00 0.10 0.20 0.30 0.40 0.50 0.60 0.70 Jaccard overlap of the observed result sets
both queries returned 10 results at least one returned fewer *

Nine of the twelve dots are hollow. Treat the horizontal positions as approximate. An incomplete capture shrinks the union and pushes a dot right, as the next section shows. Every plotted value appears in the table above.

Why the highest overlap score is misleading

B1–B2 (cek versus mengecek) has the highest observed-set Jaccard: 0.636. B1 also stopped after eight results. Here is what that does to the calculation.

B1 returned 8 URLs, B2 returned 10, and 7 were shared.
Jaccard = 7 / (8 + 10 − 7) = 7 / 11 = 0.636

Had B1 returned a full 10, with its 2 extra URLs both non-shared
(so 3 non-shared in total):
Jaccard = 7 / (10 + 10 − 7) = 7 / 13 = 0.538

0.538 is exactly the A1–A2 value.

B1–B2's apparent lead over A1–A2 rests on B1 returning two fewer results. The top-five figure, 4 of 5 shared, does not have that problem. The observed-set value does.

Nine rows have unequal or short sets. Sorting this table by Jaccard would partly sort it by parser performance.

What was in the capture

The comparison code handled both cases. Why Google returned those sets is outside this test.

Wording changed, but so did search intent

cekmengecek changes the word form but leaves the task alone. cara cekkenapa changes the task from verification to explanation.

Calling both of these “grammar changes” would muddle two different treatments. The larger study would need separate groups for intent-preserving and intent-changing pairs.

Before the next run

We would capture each query on at least three dates and in two independent sessions per date. The collector would record SERP modules and retry short blocks. Query pairs would change one linguistic feature at a time, with entity, proposition, specificity, and intent held fixed where possible.

Hypothesis boundary

The test log ends here. The collector never opened the ranking pages. The notes below describe how we currently approach page structure; this run did not test them.

Turning a query into a content outline

For writing purposes, we use this rough sequence:

query grammar → inferred intent → expected answer type → passage structure

In plain English: read the query, identify what the searcher expects to receive, then choose a page structure that delivers it. This is a writing aid, not a ranking model.

Query frameAnswer type impliedPassage structure we would try
apa itu XDefinitiondefinition → function → example
bagaimana / cara XProcedureoutcome → prerequisites → ordered steps
kenapa XExplanationmost likely cause → other causes → diagnosis → remedy
apa saja syarat XRequirementschecklist → conditions → exceptions
siapa yang wajib XEligibilityqualifying groups → exclusions → authoritative basis
X aktif atau tidakVerificationchecking method → what each status means → next action

These mappings come from content practice. This run did not test any of them.

Example: a BPJS procedure

For a procedural query, we would put the usable steps on the page:

<h1>Cara Mengecek Status BPJS Kesehatan</h1>
<p>Status kepesertaan BPJS Kesehatan dapat diperiksa melalui
aplikasi Mobile JKN.</p>

<h2>Cara mengecek melalui Mobile JKN</h2>
<ol>
  <li>Buka aplikasi Mobile JKN.</li>
  <li>Masuk menggunakan NIK atau nomor kartu dan kata sandi.</li>
  <li>Pilih menu <strong>Info Peserta</strong> di halaman utama.</li>
  <li>Periksa Status Kepesertaan: Aktif atau Tidak Aktif.</li>
</ol>

Steps verified 19 July 2026 against the BPJS Kesehatan Mobile JKN user manual. The menu is labelled Info Peserta, not "Informasi Peserta." An earlier draft used the wrong label.

The imperative sentences are there because these are instructions. We would not count imperatives or spread them through unrelated copy to satisfy a score.

A well-structured procedure can still be wrong. Device-specific instructions go stale silently, and fluent grammar can hide the error. The stamp therefore carries a date and source. A real study also needs a validity check outside the structure score.

Write passages that make sense on their own

Compare a paragraph that names its subject:

NPWP is the taxpayer identification number used in Indonesian tax administration. For Indonesian residents the 16-digit NIK now serves as the NPWP, and individuals register or activate it through the Directorate General of Taxes' Coretax system.

with one that drops the context:

It is important for many people. You can register there if required.

The first still makes sense when quoted on its own. The second does not.

NPWP/NIK and Coretax registration verified 19 July 2026 against the DJP Coretaxpedia individual registration guide and coretaxdjp.pajak.go.id.

HTML can expose that structure, but it cannot supply missing substance. We use headings for answer boundaries, ordered lists for procedures, unordered lists for requirements, tables for comparisons, and figures when a claim needs visual evidence.

What this does not prove

This run tells us nothing about ranking or citation. It did not inspect page content, record cited URLs, match citations to supporting passages, or repeat the capture to test stability.

A later citation study would record the ranking URL, the domain and URL cited in the AI answer, the supporting passage, its fit with the requested answer type, and whether the citation reappears across sessions.

How an SEO writer can use this

Before drafting, label the query as a definition, procedure, explanation, requirements list, eligibility question, or verification task. Put that answer near the top, use the matching format, keep important passages understandable on their own, cite the evidence, and date instructions that can go stale.

Those are editorial choices, not findings from this test. The report does not show that changing headings, adding steps, or rewriting sentences will improve rank.

The comparison code worked. The collector needs another pass.

Related

This is report 001 in the WebmasterNinja R&D report. The GEO scoring methodology explains how the product score is built and what it refuses to score.

Add to Chrome · free Back to overview