Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
135 changes: 135 additions & 0 deletions .ai/docs/plans/2026-09-30-indexer-heading-id-normalization.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,135 @@
# Indexer heading ID normalization implementation plan

> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.

**Goal:** Make Algolia search-result fragments match the heading IDs exposed by the rendered page.

**Architecture:** Move the existing pure heading-ID normalizer from `ak.js` into `scripts/utils/heading-id.js`. Re-export it from `ak.js` for current browser consumers, and call the same utility when the indexer reads heading IDs in `splitSections()`.

**Tech Stack:** JavaScript ES modules, Web Test Runner, Node test runner, `node-html-parser`

---

### Task 1: Share heading ID normalization with the indexer

**Files:**
- Create: `scripts/utils/heading-id.js`
- Modify: `scripts/ak.js:382-401`
- Modify: `tools/indexer/sections.js:12-61`
- Test: `test/indexer/sections.node.test.js:38-42`
- Verify: `test/scripts/ak.test.js:41-62`

- [ ] **Step 1: Write the failing indexer regression test**

Extend the existing anchor test in `test/indexer/sections.node.test.js`:

```js
it('normalizes boundary size modifiers in heading anchors', () => {
const el = main(`
<h1 id="size-xl-overview">Overview</h1><p>intro</p>
<h2 id="details-size-m">Details</h2><p>details</p>
<h3 id="size-xl-api-size-s">API</h3><p>api</p>
<h2 id="heading-size-xl-details">Middle</h2><p>middle</p>
<h2 id="sizeable-content">Similar</h2><p>similar</p>
`);

assert.deepEqual(
splitSections(el).map((section) => section.anchor),
['overview', 'details', 'api', 'heading-size-xl-details', 'sizeable-content'],
);
});
```

- [ ] **Step 2: Run the indexer section test and verify RED**

Run:

```bash
node --test test/indexer/sections.node.test.js
```

Expected: FAIL because `splitSections()` returns the original EDS IDs with boundary `size-*` modifiers.

- [ ] **Step 3: Create the shared pure utility**

Create `scripts/utils/heading-id.js`:

```js
export default function normalizeHeadingId(id) {
return id
.replace(/^size-[a-z0-9]+-/, '')
.replace(/-size-[a-z0-9]+$/, '');
}
```

The module must remain free of DOM and browser-global dependencies so both `ak.js` and the Node indexer can import it.

- [ ] **Step 4: Preserve the existing `ak.js` API**

At the top of `scripts/ak.js`, import the utility:

```js
import normalizeHeadingId from './utils/heading-id.js';
```

Replace the existing function declaration with a re-export:

```js
export { normalizeHeadingId };
```

Keep the existing `loadIcons()` call unchanged:

```js
parent.id = normalizeHeadingId(parent.id);
```

- [ ] **Step 5: Normalize IDs at the indexer boundary**

At the top of `tools/indexer/sections.js`, import the utility:

```js
import normalizeHeadingId from '../../scripts/utils/heading-id.js';
```

Normalize only heading IDs as each section is created:

```js
anchor: normalizeHeadingId(node.getAttribute('id') || ''),
```

Do not modify heading text, non-heading IDs, record construction, or source HTML.

- [ ] **Step 6: Run targeted tests and verify GREEN**

Run:

```bash
node --test test/indexer/sections.node.test.js
npm run test:file -- test/scripts/ak.test.js
```

Expected: both commands pass. The existing `ak.js` tests prove its public named export retains the same normalization behavior.

- [ ] **Step 7: Run complete relevant suites**

Run:

```bash
npm run test:indexer
npm run test:unit
npx eslint scripts/ak.js scripts/utils/heading-id.js tools/indexer/sections.js test/indexer/sections.node.test.js
git diff --check
```

Expected: all commands exit successfully. The unit suite may continue to report the repository's intentional skipped tests.

- [ ] **Step 8: Review the final diff**

Run:

```bash
git --no-pager diff -- scripts/ak.js scripts/utils/heading-id.js tools/indexer/sections.js test/indexer/sections.node.test.js
```

Confirm the diff contains only the shared utility extraction, the indexer call, and its regression coverage. Do not commit unless the user explicitly requests it.
Original file line number Diff line number Diff line change
@@ -0,0 +1,37 @@
# Indexer heading ID normalization

## Problem

EDS includes authored heading-size syntax in generated heading IDs. For example, a heading can reach the indexer with `id="size-xl-overview"`.

The browser later removes the leading or trailing `size-*` modifier through `normalizeHeadingId()`. The indexer currently copies the original source ID into Algolia records, so search results link to `#size-xl-overview` while the rendered page exposes `#overview`. Those fragment links fail.

## Scope

Normalize heading anchors while building search sections. Do not rewrite upstream EDS HTML or custom-domain HTML responses.

The normalization applies only to IDs read from `h1`, `h2`, and `h3` elements by the indexer. Preserve:

- IDs without a boundary size modifier.
- `size-*` text in the middle of an ID.
- IDs on non-heading elements.
- Empty heading IDs.

## Design

Move the pure `normalizeHeadingId(id)` function into a browser- and Node-compatible shared utility. Keep the existing `ak.js` export so current browser consumers do not change.

When `splitSections()` creates a section, pass the heading's `id` through the shared function before assigning `section.anchor`. The rest of record generation remains unchanged and builds its URL from the normalized anchor.

This keeps runtime DOM IDs and indexed anchors on one canonical normalization rule without adding HTML response parsing or duplicating the regular expressions.

## Testing

Add an indexer section test that verifies:

- `size-xl-overview` becomes `overview`.
- `overview-size-m` becomes `overview`.
- Boundary modifiers on both sides are removed.
- Middle `size-*` text and similar non-modifier text remain unchanged.

Retain the existing `ak.js` normalization tests to verify the public export still behaves the same after extraction. Run the targeted `ak.js` and indexer tests, followed by the existing unit and indexer suites.
8 changes: 3 additions & 5 deletions scripts/ak.js
Original file line number Diff line number Diff line change
Expand Up @@ -10,6 +10,8 @@
* governing permissions and limitations under the License.
*/

import normalizeHeadingId from './utils/heading-id.js';

const LOG = async (ex, el) => (await import('./utils/error.js')).default(ex, el);

export function getMetadata(name) {
Expand Down Expand Up @@ -379,11 +381,7 @@ function decorateLinks(el) {
}, []);
}

export function normalizeHeadingId(id) {
return id
.replace(/^size-[a-z0-9]+-/, '')
.replace(/-size-[a-z0-9]+$/, '');
}
export { normalizeHeadingId };

function loadIcons(el) {
const icons = [...el.querySelectorAll('span.icon')];
Expand Down
5 changes: 5 additions & 0 deletions scripts/utils/heading-id.js
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
export default function normalizeHeadingId(id) {
return id
.replace(/^size-[a-z0-9]+-/, '')
.replace(/-size-[a-z0-9]+$/, '');
}
15 changes: 15 additions & 0 deletions test/indexer/sections.node.test.js
Original file line number Diff line number Diff line change
Expand Up @@ -40,6 +40,21 @@ describe('splitSections', () => {
assert.deepEqual(splitSections(el).map((s) => s.anchor), ['one', '']);
});

it('normalizes boundary size modifiers in heading anchors', () => {
const el = main(`
<h1 id="size-xl-overview">Overview</h1><p>intro</p>
<h2 id="details-size-m">Details</h2><p>details</p>
<h3 id="size-xl-api-size-s">API</h3><p>api</p>
<h2 id="heading-size-xl-details">Middle</h2><p>middle</p>
<h2 id="sizeable-content">Similar</h2><p>similar</p>
`);

assert.deepEqual(
splitSections(el).map((section) => section.anchor),
['overview', 'details', 'api', 'heading-size-xl-details', 'sizeable-content'],
);
});

it('tracks hierarchy across levels and resets deeper levels', () => {
const el = main(`
<h1 id="p">Page</h1><p>a</p>
Expand Down
4 changes: 3 additions & 1 deletion tools/indexer/sections.js
Original file line number Diff line number Diff line change
Expand Up @@ -9,6 +9,8 @@
* aggregate text, so nesting cannot count a passage twice.
*/

import normalizeHeadingId from '../../scripts/utils/heading-id.js';

const HEADING_TAGS = new Set(['h1', 'h2', 'h3']);
const MAX_CONTENT_LENGTH = 8000;
const TEXT_NODE = 3;
Expand Down Expand Up @@ -57,7 +59,7 @@ export function splitSections(main) {
current = {
heading: normalize(node.text),
level: Number(tag[1]),
anchor: node.getAttribute('id') || '',
anchor: normalizeHeadingId(node.getAttribute('id') || ''),
parts: [],
};
// Do not descend: the heading's own text is not body content.
Expand Down
Loading