# PWA 的 SEO 与结构化数据

> JSON-LD 结构化数据与可抓取的 HTML 如何让搜索引擎理解 PWA 并获得富媒体搜索结果，以及如何用富媒体搜索结果测试验证。

完成本指南后，PWA 的每个公开页面都带有一段 Google 富媒体搜索结果测试能无错解析的 `WebApplication` JSON-LD，该标记与可见页面出自同一份数据，并且构建时检查会在标记无法解析时失败。以网站形式分发的 PWA 与其他站点一样被索引；MDN 的可安装性指南称之为 "making it discoverable through web search"。manifest 与 Service Worker 对索引没有贡献，所以本指南讨论的是爬虫看到的 HTML。

你需要把主要内容渲染为 HTML 的页面（服务端渲染或预渲染）、每页一个公开 URL，以及修改 `<head>` 模板的权限。

## 选择 schema.org 类型与属性

schema.org 在 `Thing > CreativeWork > SoftwareApplication > WebApplication` 下定义 `WebApplication`。它自有的属性是 `browserRequirements`，描述为 "Specifies browser requirements in human-readable text. For example, 'requires HTML5 support'"。有用的属性来自 `SoftwareApplication`：`applicationCategory`、`operatingSystem`、`screenshot`、`softwareVersion`、`featureList`，以及（继承自 `CreativeWork`）表示免费或付费的 `offers`。取值直接用 manifest 里已有的数据，这样 `name`、`description` 与截图 URL 就不会偏离安装界面显示的内容。

## 由 manifest 数据生成 JSON-LD

Google Search Central 推荐 JSON-LD，因为它比 Microdata 或 RDFa "less prone to user errors"，并接受它出现在 "a `<script>` tag in the `<head>` and `<body>` elements"。在页面模板中，每页序列化一个对象：

```html
<script type="application/ld+json">
{
  "@context": "https://schema.org",
  "@type": "WebApplication",
  "name": "Field Notes",
  "url": "https://notes.example/",
  "description": "Offline-first notes that sync when you are back online.",
  "applicationCategory": "ProductivityApplication",
  "operatingSystem": "Any",
  "browserRequirements": "Requires JavaScript and a browser with service worker support.",
  "screenshot": "https://notes.example/screenshots/desktop-wide.png",
  "offers": { "@type": "Offer", "price": "0", "priceCurrency": "USD" }
}
</script>
```

Google 的通用准则适用于每个属性："don't add structured data about information that is not visible to the user, even if the information is accurate"，以及 "Don't create blank or empty pages just to hold structured data"。页面上没有显示的 `screenshot`，或与可见文案不一致的 `description`，即使事实正确也违反准则。

## 让内容在没有 Service Worker 时也可抓取

爬虫不会运行你的 Service Worker，所以内容要等缓存的应用壳启动后才出现的页面，会被当作空壳索引。把文章、商品或笔记正文渲染进 HTML 响应，再让 Service Worker 缓存同一份响应。Google 表示它 "can read JSON-LD data when it is dynamically injected into the page's contents"，因此标记本身用客户端注入可以被接受，但写在响应里的静态块能去掉渲染期间对 JavaScript 执行的依赖。

## 用富媒体搜索结果测试与构建检查验证

打开富媒体搜索结果测试，在 **Enter a URL to test** 中粘贴公开 URL，或切到 **Code** 粘贴 HTML，然后运行。结果会列出检测到的条目以及解析或必填属性错误。在构建中加一道检查，解析每段输出的标记，让模板 bug 在部署前失败，而不是在 Search Console 里暴露：

```js
import { readFile } from 'node:fs/promises';

export async function assertJsonLd(htmlPath) {
  const html = await readFile(htmlPath, 'utf8');
  const blocks = [...html.matchAll(/<script type="application\/ld\+json">([\s\S]*?)<\/script>/g)];
  if (blocks.length === 0) throw new Error(`${htmlPath}: no JSON-LD block`);
  for (const [, json] of blocks) {
    const data = JSON.parse(json); // 模板损坏时抛出异常
    if (data['@type'] !== 'WebApplication') throw new Error(`${htmlPath}: unexpected @type ${data['@type']}`);
  }
}
```

`JSON.parse` 能发现格式错误的输出，但不检查 schema.org 词汇，后者正是富媒体搜索结果测试的工作，所以两者都要跑。

:::observed
富媒体搜索结果测试（search.google.com/test/rich-results，英文界面，2026-10-03 读取）以标题 "Does your page support rich results?" 开头，提供 **Enter a URL to test** 与 **Code** 两种输入，并可在 "Google Inspection Tool smartphone" 与 "Google Inspection Tool desktop" 之间选择设备。URL 模式会像 Googlebot 一样抓取线上页面，所以只有在无需登录即可访问时，才能用它测试预发布 URL。
:::

## 另请参阅

- [让它可安装](/zh/guides/installable/)
- [`screenshots` manifest 成员](/zh/reference/manifest/screenshots/)
- [`description` manifest 成员](/zh/reference/manifest/description/)
- [Introduction to structured data markup in Google Search](https://developers.google.com/search/docs/appearance/structured-data/intro-structured-data)（developers.google.com）
- [WebApplication](https://schema.org/WebApplication)（schema.org）