AIクローラーの本物と偽物を見分ける実装。OpenAI・Anthropic・Perplexity・Googleが公開しているIPアドレスの範囲と照らし合わせるモジュールを、Cloudflare Pages Functions向けに書き、実際の範囲と偽のIPでテストした
GPTBotなどを
この
この
各社が公開しているIPアドレスの一覧
2026年10月5日の
| 会社 | 一覧の |
対象 | 取得した |
一覧の |
|---|---|---|---|---|
https://developers.google.com/static/crawling/ipranges/common-crawlers.json |
Googlebotなど |
317件 |
2026-10-02 | |
| OpenAI | https://openai.com/gptbot.json |
GPTBot | 18件 | 2026-09-22 |
| OpenAI | https://openai.com/searchbot.json |
OAI-SearchBot | 39件 | 2026-01-02 |
| OpenAI | https://openai.com/chatgpt-user.json |
ChatGPT-User | 230件 | 2026-09-25 |
| Anthropic | https://claude.com/crawling/bots.json |
Anthropicの |
28件 | 2026-10-02 |
| Perplexity | https://www.perplexity.ai/perplexitybot.json |
PerplexityBot | 8件 | 2025-02-07 |
| Perplexity | https://www.perplexity.ai/perplexity-user.json |
Perplexity-User | 4件 | 2025-10-17 |
件数はprefixes のipv4Prefix か ipv6Prefix を20.171.206.0/24 など)で
取りに
- Googleの
一覧の 以前のURLが 変わっていた。 https://developers.google.com/static/search/apis/ipranges/googlebot.jsonを取ると、 301で common-crawlers.jsonに転送されました。 Googleの 今の 説明ページも、 common-crawlers.jsonを案内しています。 古い URLを 書いたままの スクリプトは、 転送を たどらないと 取れません - Perplexityの
一覧は、 説明ページには説明ページの URLと 実際の 置き場所が 違う。 https://www.perplexity.com/perplexitybot.jsonと書かれていますが、 取ると 302で https://www.perplexity.ai/perplexitybot.jsonに転送されました - OpenAIには、
もう 広告と1つ 一覧が ある。 して 出すページの 安全を 確かめる OAI-AdsBot の https://openai.com/adsbot.jsonです。AI検索や 学習とは 目的が 違うので、 この 記事の モジュールには 入れていません
AnthropicのClaudeBot など)だけが
Googleに
名乗りの判定と、範囲の照合は分けて考える
判定は
| User-Agent | IPアドレス | 判定 | 記録するか |
|---|---|---|---|
| クローラーを |
(見ない) | null |
しない |
| 名乗っている | 名乗った |
verified: true(本物) |
する |
| 名乗っている | 名乗った |
verified: false(名乗りだけ) |
する |
User-Agentを
また、
一覧を1つのファイルにまとめるスクリプト
一覧は、
どれか
// 各社が公開しているクローラーのIP範囲(JSON)を取ってきて、1つのファイル bot-ranges.json にまとめる
// 使い方: node fetch-ranges.mjs (ビルドの前に流し、できたファイルを Functions から import する)
import { writeFile } from "node:fs/promises";
const SOURCES = {
google: "https://developers.google.com/static/crawling/ipranges/common-crawlers.json",
gptbot: "https://openai.com/gptbot.json",
"oai-searchbot": "https://openai.com/searchbot.json",
"chatgpt-user": "https://openai.com/chatgpt-user.json",
anthropic: "https://claude.com/crawling/bots.json",
perplexitybot: "https://www.perplexity.ai/perplexitybot.json",
"perplexity-user": "https://www.perplexity.ai/perplexity-user.json",
};
const out = {};
for (const [key, url] of Object.entries(SOURCES)) {
const res = await fetch(url);
if (!res.ok) throw new Error(`${key}: HTTP ${res.status}(前回のファイルを残すため、書き出さずに止める)`);
const json = await res.json();
const prefixes = (json.prefixes ?? []).map((p) => p.ipv4Prefix ?? p.ipv6Prefix).filter(Boolean);
if (prefixes.length === 0) throw new Error(`${key}: 範囲が0件(形式が変わった可能性)`);
out[key] = { source: url, creationTime: json.creationTime ?? "", prefixes };
console.log(`${key.padEnd(16)} ${String(prefixes.length).padStart(4)}件 ${json.creationTime ?? ""}`);
}
await writeFile(new URL("./bot-ranges.json", import.meta.url), JSON.stringify(out));google 317件 2026-10-02T14:46:11.000000
gptbot 18件 2026-09-22T02:00:07.000000
oai-searchbot 39件 2026-01-02T11:00:00.000000
chatgpt-user 230件 2026-09-25T18:04:57.257335
anthropic 28件 2026-10-02T00:00:00Z
perplexitybot 8件 2025-02-07T16:56:00.000000
perplexity-user 4件 2025-10-17T10:17:00.000000できた
照合するモジュール
IPアドレスと
// User-Agent で名乗ったクローラーが、その会社の公開しているIP範囲から来ているかを確かめる
// Cloudflare Pages Functions(Workers)でも Node でも、そのまま動く(依存なし)
// 名乗り(User-Agent)と、照合に使う範囲の対応。上から順に見る
export const BOTS = [
{ name: "OAI-SearchBot", ua: /oai-searchbot/i, ranges: ["oai-searchbot"] },
{ name: "ChatGPT-User", ua: /chatgpt-user/i, ranges: ["chatgpt-user"] },
{ name: "GPTBot", ua: /gptbot/i, ranges: ["gptbot"] },
// Anthropic は ClaudeBot・Claude-User・Claude-SearchBot の3つで1つのリスト
{ name: "Claude-SearchBot", ua: /claude-searchbot/i, ranges: ["anthropic"] },
{ name: "Claude-User", ua: /claude-user/i, ranges: ["anthropic"] },
{ name: "ClaudeBot", ua: /claudebot/i, ranges: ["anthropic"] },
{ name: "Perplexity-User", ua: /perplexity-user/i, ranges: ["perplexity-user"] },
{ name: "PerplexityBot", ua: /perplexitybot/i, ranges: ["perplexitybot"] },
{ name: "Googlebot", ua: /googlebot/i, ranges: ["google"] },
];
// "1.2.3.4" や "2001:db8::1" を [版, BigInt] にする。読めなければ null
export function parseIp(ip) {
if (typeof ip !== "string") return null;
ip = ip.trim();
// IPv4 を IPv6 の形で書いたもの(::ffff:1.2.3.4)は IPv4 として扱う
const mapped = /^::ffff:(\d+\.\d+\.\d+\.\d+)$/i.exec(ip);
if (mapped) ip = mapped[1];
if (/^\d+\.\d+\.\d+\.\d+$/.test(ip)) {
const p = ip.split(".").map(Number);
if (p.some((n) => n > 255)) return null;
return [4, p.reduce((acc, n) => (acc << 8n) + BigInt(n), 0n)];
}
if (!ip.includes(":") || !/^[0-9a-f:]+$/i.test(ip)) return null;
const halves = ip.split("::");
if (halves.length > 2) return null;
const head = halves[0] ? halves[0].split(":") : [];
const tail = halves.length === 2 && halves[1] ? halves[1].split(":") : [];
const fill = halves.length === 2 ? 8 - head.length - tail.length : 0;
if (fill < 0 || (halves.length === 1 && head.length !== 8)) return null;
const groups = [...head, ...Array(fill).fill("0"), ...tail];
if (groups.some((g) => g.length === 0 || g.length > 4)) return null;
return [6, groups.reduce((acc, g) => (acc << 16n) + BigInt(parseInt(g, 16)), 0n)];
}
// "20.171.206.0/24" を { v, net, mask } にする
function parseCidr(cidr) {
const [addr, bitsStr] = cidr.split("/");
const parsed = parseIp(addr);
if (!parsed) return null;
const [v, n] = parsed;
const total = v === 4 ? 32 : 128;
const bits = bitsStr === undefined ? total : Number(bitsStr);
if (!Number.isInteger(bits) || bits < 0 || bits > total) return null;
const mask = bits === 0 ? 0n : ((1n << BigInt(bits)) - 1n) << BigInt(total - bits);
return { v, net: n & mask, mask };
}
// bot-ranges.json の形 { key: { prefixes: [...] } } を、照合しやすい形に一度だけ変換する
export function compileRanges(json) {
const table = {};
for (const [key, { prefixes }] of Object.entries(json)) {
table[key] = prefixes.map(parseCidr).filter(Boolean);
}
return table;
}
export function ipInRanges(ip, list) {
const parsed = parseIp(ip);
if (!parsed || !list) return false;
const [v, n] = parsed;
return list.some((r) => r.v === v && (n & r.mask) === r.net);
}
// 戻り値: null(クローラーを名乗っていない)か、{ bot, verified }
// verified: true=公開範囲の中から来た、false=名乗っているのに範囲の外から来た
export function classify(ua, ip, table) {
const hit = BOTS.find((b) => b.ua.test(ua || ""));
if (!hit) return null;
const verified = hit.ranges.some((key) => ipInRanges(ip, table[key]));
return { bot: hit.name, verified };
}作りで
- 一覧の
変換は 1回だけ。 compileRangesで、CIDRの 文字列を 数と マスクに 直した ものを 作って おき、 リクエストごとには 比べるだけに しています ::ffff:1.2.3.4はIPv4と IPv4のして 扱う。 アドレスを IPv6の 形で 書いた ものです。 そのまま 比べると、 IPv4の 一覧に 当たらなくなります - 読めない
値は 空の偽物に する。 文字列や 999.1.1.1のような値でも、 例外は 投げずに verified: falseを返します - 名前は
上から 最初に順に 見る。 当たった 名前を 使います。 いまの 各社の 名前どうしでは、 一方が 他方の 一部に なっている ものは ありませんが、 名前を 足すときは、 長く 具体的な 名前を 先に 置くように しています
実際の範囲と偽のIPアドレスで、手元のNodeでテストした
テストは、node:test で
- 7つの
一覧の 644件すべてに ついて、 範囲の 先頭と 末尾の アドレスが 「本物」に なる - 名乗りが
本物でも、 文書の 例に 使う ために 予約された アドレス (IPv4の 192.0.2.0/24・198.51.100.0/24・203.0.113.0/24、IPv6の 2001:db8::/32)から来たら 「偽物」に なる - 別の
会社の 範囲から 来たら 「偽物」に なる - 範囲の
すぐ 外 (1つ前と 1つ 後)は 「偽物」に なる - クローラーを
名乗っていなければ nullになる - 壊れた
入力でも 例外を 投げない
// 実際に取ってきた bot-ranges.json と、偽物のIP(文書用に予約されたアドレス)で確かめる
import { test } from "node:test";
import assert from "node:assert/strict";
import { readFile } from "node:fs/promises";
import { parseIp, compileRanges, ipInRanges, classify } from "./bot-verify.js";
const raw = JSON.parse(await readFile(new URL("./bot-ranges.json", import.meta.url), "utf8"));
const table = compileRanges(raw);
// 範囲の先頭と末尾のアドレスを文字列で作る(末尾=ホスト部をすべて1にしたもの)
function edges(cidr) {
const [addr, bits] = cidr.split("/");
const [v, n] = parseIp(addr);
const total = v === 4 ? 32 : 128;
const last = n | ((1n << BigInt(total - Number(bits))) - 1n);
const fmt = (x) => v === 4
? [24n, 16n, 8n, 0n].map((s) => String((x >> s) & 255n)).join(".")
: Array.from({ length: 8 }, (_, i) => ((x >> BigInt(112 - 16 * i)) & 0xffffn).toString(16)).join(":");
return [fmt(n), fmt(last)];
}
// UA の全文は、各社のページに載っているもの
const UA = {
GPTBot: "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot",
"OAI-SearchBot": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36; compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot",
"ChatGPT-User": "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; ChatGPT-User/1.0; +https://openai.com/bot",
// Anthropic のページには UA の全文が載っていないので、名前だけを入れた仮の文字列
ClaudeBot: "Mozilla/5.0 (compatible; ClaudeBot)",
PerplexityBot: "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot)",
"Perplexity-User": "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Perplexity-User/1.0; +https://perplexity.ai/perplexity-user)",
Googlebot: "Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)",
};
const KEY = { GPTBot: "gptbot", "OAI-SearchBot": "oai-searchbot", "ChatGPT-User": "chatgpt-user", ClaudeBot: "anthropic", PerplexityBot: "perplexitybot", "Perplexity-User": "perplexity-user", Googlebot: "google" };
// 公開されているすべての範囲について、先頭と末尾のアドレスが「本物」になること
for (const [bot, key] of Object.entries(KEY)) {
test(`${bot}: 公開範囲 ${raw[key].prefixes.length} 件の先頭と末尾がすべて本物になる`, () => {
for (const cidr of raw[key].prefixes) {
for (const ip of edges(cidr)) {
assert.deepEqual(classify(UA[bot], ip, table), { bot, verified: true }, `${cidr} の ${ip}`);
}
}
});
}
test("名乗りは本物でも、文書用のアドレスから来たら偽物になる", () => {
for (const ip of ["192.0.2.10", "198.51.100.7", "203.0.113.45", "2001:db8::1"]) {
for (const bot of Object.keys(UA)) assert.equal(classify(UA[bot], ip, table).verified, false, `${bot} ${ip}`);
}
});
test("別の会社の範囲から来たら偽物になる(GPTBotを名乗ってAnthropicの範囲から)", () => {
const [anthropicIp] = edges(raw.anthropic.prefixes[0]);
assert.equal(classify(UA.GPTBot, anthropicIp, table).verified, false);
const [gptIp] = edges(raw.gptbot.prefixes[0]);
assert.equal(classify(UA.ClaudeBot, gptIp, table).verified, false);
});
test("範囲のすぐ外は偽物になる(gptbot.json の 20.125.66.80/28 の前後)", () => {
assert.equal(ipInRanges("20.125.66.79", table.gptbot), false);
assert.equal(ipInRanges("20.125.66.80", table.gptbot), true);
assert.equal(ipInRanges("20.125.66.95", table.gptbot), true);
assert.equal(ipInRanges("20.125.66.96", table.gptbot), false);
});
test("IPv4 を IPv6 の形(::ffff:)で書いても同じ結果になる", () => {
const [ip] = edges(raw.gptbot.prefixes[0]);
assert.equal(classify(UA.GPTBot, `::ffff:${ip}`, table).verified, true);
});
test("クローラーを名乗っていなければ null(記録しない)", () => {
const chrome = "Mozilla/5.0 (iPhone; CPU iPhone OS 18_0 like Mac OS X) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/18.0 Mobile/15E148 Safari/604.1";
assert.equal(classify(chrome, "203.0.113.45", table), null);
assert.equal(classify("", "203.0.113.45", table), null);
});
test("壊れた入力でも例外を投げず、偽物として扱う", () => {
for (const ip of ["", "999.1.1.1", "1.2.3", "2001:db8:::1", "hello", null, undefined]) {
assert.equal(classify(UA.GPTBot, ip, table).verified, false, String(ip));
}
});✔ GPTBot: 公開範囲 18 件の先頭と末尾がすべて本物になる (1.3805ms)
✔ OAI-SearchBot: 公開範囲 39 件の先頭と末尾がすべて本物になる (0.813792ms)
✔ ChatGPT-User: 公開範囲 230 件の先頭と末尾がすべて本物になる (2.566583ms)
✔ ClaudeBot: 公開範囲 28 件の先頭と末尾がすべて本物になる (0.279791ms)
✔ PerplexityBot: 公開範囲 8 件の先頭と末尾がすべて本物になる (0.114166ms)
✔ Perplexity-User: 公開範囲 4 件の先頭と末尾がすべて本物になる (0.047208ms)
✔ Googlebot: 公開範囲 317 件の先頭と末尾がすべて本物になる (3.125875ms)
✔ 名乗りは本物でも、文書用のアドレスから来たら偽物になる (0.28775ms)
✔ 別の会社の範囲から来たら偽物になる(GPTBotを名乗ってAnthropicの範囲から) (0.244833ms)
✔ 範囲のすぐ外は偽物になる(gptbot.json の 20.125.66.80/28 の前後) (0.255834ms)
✔ IPv4 を IPv6 の形(::ffff:)で書いても同じ結果になる (0.099667ms)
✔ クローラーを名乗っていなければ null(記録しない) (0.038792ms)
✔ 壊れた入力でも例外を投げず、偽物として扱う (0.056375ms)
ℹ tests 13
ℹ suites 0
ℹ pass 13
ℹ fail 0
ℹ cancelled 0
ℹ skipped 0
ℹ todo 0
ℹ duration_ms 67.993167テストがbot-verify.js のn === r.net)に
速さも
// いちばん時間のかかる形(Googlebot を名乗って範囲の外から来る=317件をすべて見る)を10万回
import { readFile } from "node:fs/promises";
import { compileRanges, classify } from "./bot-verify.js";
const table = compileRanges(JSON.parse(await readFile(new URL("./bot-ranges.json", import.meta.url), "utf8")));
const ua = "Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)";
for (const ip of ["203.0.113.45", "2001:db8::1"]) {
const N = 100_000, t0 = performance.now();
for (let i = 0; i < N; i++) classify(ua, ip, table);
const ms = performance.now() - t0;
console.log(`${ip.padEnd(14)} 1回あたり ${((ms / N) * 1000).toFixed(1)} マイクロ秒`);
}203.0.113.45 1回あたり 0.9 マイクロ秒
2001:db8::1 1回あたり 1.7 マイクロ秒Workersの
Googlebotは、DNSの方法でも同じ結果になった
Googleが66.249.66.1 にhost コマンドを
$ host 66.249.66.1
1.66.249.66.in-addr.arpa domain name pointer crawl-66-249-66-1.googlebot.com.
$ host crawl-66-249-66-1.googlebot.com
crawl-66-249-66-1.googlebot.com has address 66.249.66.1
$ host 203.0.113.45
Host 45.113.0.203.in-addr.arpa. not found: 3(NXDOMAIN)逆引きのgooglebot.com で66.249.66.1 は203.0.113.45 は、
Pages Functionsで読み込めるかを、wrangler pages devで確かめた
最後に、
// 手元の確認用:クローラーを名乗ったリクエストだけ、判定結果をログに出す(応答は変えない)
import { compileRanges, classify } from "./bot-verify.js";
import ranges from "./bot-ranges.json";
const table = compileRanges(ranges); // 起動時に1回だけ変換する
export async function onRequest({ request, next }) {
const res = await next();
try {
const result = classify(request.headers.get("user-agent"), request.headers.get("cf-connecting-ip"), table);
if (result) {
console.log(JSON.stringify({ ...result, path: new URL(request.url).pathname, ip: request.headers.get("cf-connecting-ip") }));
}
} catch (_) {
// 判定の失敗は応答に関係させない
}
return res;
}npx wrangler pages dev public --port 8837 --compatibility-date 2026-10-01
# 別の端末から
G="Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot"
curl -s -o /dev/null -w "%{http_code}\n" -A "$G" localhost:8837/
curl -s -o /dev/null -w "%{http_code}\n" -A "$G" -H "CF-Connecting-IP: 20.171.206.10" localhost:8837/
curl -s -o /dev/null -w "%{http_code}\n" -A "Mozilla/5.0 (iPhone)" localhost:8837/[wrangler:info] Ready on http://localhost:8837
[wrangler:info] GET / 200 OK (15ms)
{"bot":"GPTBot","verified":false,"path":"/","ip":"::1"}
[wrangler:info] GET / 200 OK (3ms)
{"bot":"GPTBot","verified":true,"path":"/","ip":"20.171.206.10"}
[wrangler:info] GET / 200 OK (2ms)
[wrangler:info] GET / 200 OK (2ms)最初の GET / 200 は、import ranges from "./bot-ranges.json" の::1(自分のCF-Connecting-IP に
CF-Connecting-IP は、
記録するときの、個人情報の扱い
名乗りだけの
- クローラーを
名乗った 名乗っていないリクエストだけを 記録する。 人の アクセスは、 判定にも 記録にも 進みません - IPアドレスは
丸めて この残す。 ブログの 今の 記録の 処理は、 IPv4は /24、IPv6は /48に丸めてから 書く 作りです。 判定は リクエストを 受けた その 場で、 丸める 前の アドレスで 行い、 残すのは 判定の 結果 (本物か 名乗りだけか)と、 丸めた アドレスだけに します - 保存期間を
決めて このおく。 ブログの 記録の 書き先の Workers Analytics Engineは、 Cloudflareの 説明では、 書いた データを 3か 月保存します。 それより 長く 残したい 集計は、 ボット名と 判定ごとの 件数のように、 個人に 結びつかない 形に してから 別に 残します
本番に入れるときに決めること
この
- 一覧を
更新する 一覧の時機。 JSONには 作成日時 (creationTime)が 入っていて、 取得した ときは 2025年2月7日から 2026年10月2日まで 幅が ありました。 Perplexityの 説明にも、 公式の 一覧の 最新の ものを 使うよう 書かれています。 一覧が 更新されたのに 古い ものを 使い続けると、 新しい アドレスから 来た 本物を 偽物と 判定します。 ビルドの たびに fetch-ranges.mjsを流し、 一覧の 作成日時も 記録に 残す形を 考えています - 記録に
足す項目。 今の記録に、 判定の 結果 ( true・false)と、照らした 一覧の 作成日時を 足します - 偽物と
判定した すぐにアクセスを どうするか。 止めるのではなく、 数週間記録して、 件数と 中身を 見てから 決めます。 Anthropicの ヘルプには、 IPアドレスで 止める 方法は、 robots.txtを 読めなく するので、 確実な 拒否の 方法には ならないと 書かれています。 Anthropicの ヘルプが 止め方と して 案内しているのは、 robots.txtでの 指定です。 robots.txtの 書き方と Cloudflare側の 設定の 確かめ方は、AIクローラーの robots.txt設定例の に記事 書きました
Googleのcommon-crawlers.json だけで
このモジュールでは分からないこと
IPアドレスで
- 一覧に
無い 新しい アドレスから 来た 本物 (一覧が 更新される 前の 時間差) - 名乗っていない
クローラー (User-Agentを 人の ブラウザに 似せた もの) - 本物の
クローラーが、 その ページを 何に 使ったか (学習か、 検索か、 利用者の 依頼か)。 これは 名前で 分かれているので、 名前ごとの 各社の 説明を 見ます
会社の
株式会社bundlyzeでは、
動作確認した環境
確認日は
- macOS 26.6.2、
Node.js 22.23.2 ( node:test、組み込みの fetch) - wrangler 4.147.0
( pages dev、互換性の 日付は 2026-10-01) - curl 8.7.1、
hostコマンド(macOS 付属) - 一覧は、
上の 表の 7つの URLから 2026年10月5日に 取得した もの
確かめていない
- Cloudflareの
本番 (Pages Functions)での 動作と 速さ。 この モジュールは 本番に 入れていません - 本番での
CF-Connecting-IPの扱い (手元の wranglerでの 値の 入り方だけを 確かめました) - この
ブログに 実際に 来た クローラーの うち、 名乗りだけの ものが どの くらい あるか (本番に 入れてから 数えます) - IPv4を
埋め込んだ IPv6の 書き方の うち、 ::ffff:1.2.3.4以外の形 (モジュールは 読めない 値と して 偽物にします)
参照した公式ドキュメント(確認日:2026年10月5日)
- OpenAIの
クローラー (OAI-SearchBot・OAI-AdsBot・GPTBot・ChatGPT-User)の 目的、 User-Agent、 IPアドレスの 一覧の 場所 (searchbot.json・adsbot.json・gptbot.json・chatgpt-user.json)、 ChatGPT-Userは 利用者の 操作で 動くので robots.txtが 適用されない ことが ある こと → OpenAI 「Overview of OpenAI Crawlers」 - Anthropicの
クローラー (ClaudeBot・Claude-User・Claude-SearchBot)の 目的、 送信元の IPアドレスが bots.json の 一覧に あれば Anthropicから 来た クローラーである こと、 IPアドレスで 止める 方法は 確実な 拒否に ならない こと → Anthropic 「Does Anthropic crawl data from the web, and how can site owners block the crawler?」 - Perplexityの
クローラー (PerplexityBot・Perplexity-User)の 目的、 User-Agent、 IPアドレスの 一覧の 場所、 User-Agentと IPアドレスの 確認を 組み合わせるよう 勧めている こと、 公式の 一覧の 最新の ものを 使う こと → Perplexity 「Perplexity Crawlers」 - Googleの
クローラーの 確かめ方 (逆引きと 正引きの DNSで、 googlebot.com・google.com・googleusercontent.com の いずれかである ことを 確かめる 手順、 IPアドレスの 一覧との 照合、 一覧の 種類と common-crawlers.json が Googlebotなど 一般的な クローラーの 一覧である こと、 一覧に 無い IPアドレスから 来る ことが ある こと、 JSONの IPアドレスは CIDRの 形である こと) → Google 「Google からの リクエストを 」確認する - Googlebotの
User-Agentの 文字列 → Google 「Google の 一般的な 」クローラー CF-Connecting-IPがCloudflareに 接続してきた クライアントの IPアドレスを 伝える こと → Cloudflare 「Cloudflare HTTP headers」 - Workers Analytics Engineに
書いた データの 保存期間が 3か月である こと → Cloudflare 「Analytics Engine の Limits 」 - 文書の
例に 使う ために 予約された IPv4の アドレス (192.0.2.0/24・198.51.100.0/24・203.0.113.0/24) → RFC 5737 - 文書の
例に 使う ために 予約された IPv6の アドレス (2001:db8::/32) → RFC 3849
よくある質問
User-Agentに「GPTBot」と書いてあれば、OpenAIのクローラーだと考えてよいですか?
考えない
AnthropicのClaudeBotは、IPアドレスの一覧を公開していますか?
公開しています。
Googlebotは、IPアドレスの一覧とDNSのどちらで確かめればよいですか?
Googleは
名乗りが偽物のアクセスは、すぐにブロックしたほうがよいですか?
まずは