yum upgrades for production use, this is the repository for you.
Active subscription is required.
Technical Briefing: Keeping OpenAI Crawlers Unblocked — IP Feeds, NGINX Rate-Limit Bypass, and fail2ban Exemptions
The Problem
OpenAI publishes three crawler IP feeds — GPTBot, ChatGPT-User, and OAI-SearchBot — as JSON, each with a creationTime field and a prefixes array of ipv4Prefix entries. These ranges live in Microsoft Azure space and rotate, so any hard-coded list goes stale quickly.
The practical failure mode is silent: robots.txt alone is insufficient. Rate limiting (limit_req/limit_conn), allow/deny access rules, and fail2ban all act independently of robots.txt. Each can block OpenAI crawlers without any indication in the robots file — and blocked crawlers mean content drops out of ChatGPT Search results.
The Fix
The approach has three layers: bypass rate limits for OpenAI IPs, extend access-control allowlists, and exempt those IPs from fail2ban — all fed from a single auto-updating source of truth.
NGINX Rate-Limit Bypass: Two-Stage geo + map
The pattern chains a geo block and a map block:
- The
geoblock tags OpenAI IPs with""and everyone else with0. - The
mapblock converts that into either an empty string (bypass) or$binary_remote_addr(normal per-IP key). limit_req_zoneskips any request whose key is empty — so OpenAI crawlers are never counted against the limit.
The example configuration uses zone=public:10m rate=20r/s with burst=40 nodelay.
Access-Control Layer
For locked-down locations, extend allow/deny blocks with:
include /etc/nginx/iplist/openai.allow.conf;
Caveat: allow/deny only runs in the access phase. A bare return 200 short-circuits earlier and skips the check entirely.
Package-Based Delivery
The GetPageSpeed extras repo ships nginx-iplist-openai, which drops three formats into /etc/nginx/iplist/:
openai.geo.confopenai.allow.confopenai.nolimit.conf
plus a plain CIDR list at /usr/share/trusted-lists/plain/openai.txt. Updates flow through dnf update.
fail2ban Integration
The companion package fail2ban-trusted-lists-helper installs a [DEFAULT] ignorecommand hook that runs grepcidr -f against every /usr/share/trusted-lists/plain/*.txt. Because it sits in [DEFAULT], every jail inherits it. If the candidate IP matches, the script exits 0 and fail2ban skips the ban.
Verify with:
fail2ban-client get sshd ignorecommand
Composability
Installing additional list packages (Stripe, PayPal, Bingbot, etc.) automatically extends the fail2ban exemption set — no extra configuration needed.
Verification Steps
- Debug location echoing
$openai_ipand$limit_keyconfirms the geo/map wiring. - End-to-end smoke tests:
curlwith a GPTBot user-agent, expecting HTTP 200hey -n 200 -c 20load test, expecting no 429s- Running the ignore-check script directly
robots.txt Minimum
Use separate User-agent blocks for GPTBot, ChatGPT-User, and OAI-SearchBot with Allow: /. Paths not wanted for training can be denied under GPTBot while keeping the other two open.
The Payoff
The full setup is two packages plus a robots.txt edit. One source of truth auto-updates via dnf, so new OpenAI prefixes are trusted on the next update cycle — no manual list maintenance, no silent crawler blocks.
Separate Note
GetPageSpeed Amplify runs scheduled gixy scans across hosts and ties findings to live NGINX runtime metrics. It is drop-in compatible with the deprecated nginx-amplify-agent, which reaches EOL in January 2026.
Read the full article: Whitelist OpenAI IP Ranges in NGINX and fail2ban
