Field Notes · 6 MIN

How do you run AI products on Cloudflare without token sprawl?

Field notes on hosting AI product frontends on Cloudflare Pages and Workers, scoping API access once, and the WAF rule that hid a site from AI answers.

By NactorePublished 4 Sep 2026All articles

You run them by keeping the frontend on static hosting, putting small server logic in a Worker only where you need it, and settling API access once instead of minting a new token for every task. The two mistakes that cost us time were narrow tokens that could not do the job and a firewall rule that quietly blocked AI crawlers. This post covers the setup, the access model, and the checks we now run.

Key takeaways
  • Static frontends on Cloudflare Pages, deployed from CI, keep AI product launches cheap and fast to repeat.
  • Use the official Cloudflare MCP server for ops where possible, and keep one broad-scoped token for headless runs, instead of a new narrow token per task.
  • Browser OAuth through the command line tool carried read-only zone access in our setup, so it could not edit firewall rules.
  • A web application firewall can return errors to AI crawlers while serving Googlebot normally. Test with each crawler, not just a browser.
  • Nactore is an AI-native software engineering partner, and we run the infrastructure behind the AI products we ship.

What do we host on Cloudflare?

Three kinds of things. Single-page frontends go to Cloudflare Pages, with a custom domain per product. A small Worker handles server-side logic that does not justify a full backend, such as a contact form. DNS and the firewall sit in front of everything.

Cloudflare describes Workers as a serverless platform for building and deploying apps across its network without managing infrastructure. For a product team, that means small pieces of logic can live next to the static site instead of on another service.

PieceWhere it runsWhy
Product frontendPagesStatic, cached, deploys on push
Form or small API glueWorkerNo server to maintain
Heavy domain logicSeparate backendNeeds a database and long-running work
DNS, firewall rulesCloudflare zoneOne place to control the edge

How do you deploy a frontend from CI?

Cloudflare documents deploying Pages from GitHub Actions using its wrangler action with an API token and account ID stored as repository secrets. We use that pattern. A push to the main branch builds the app and deploys it. Pull requests run a build-only check.

The gain is that nobody needs a local login to ship. Merge to main and the site updates. One detail to remember is that the build runs on the CI runner, so build-time variables belong in the workflow, not the host's dashboard. We cover that in one backend, many products.

Why did we stop creating a token per task?

Each time we needed something new, we minted a new narrow token. That produced a trail of tokens with overlapping scopes and no clear owner. Worse, the original token only carried DNS edit, Pages edit, and zone read. It could not even read the zone's rulesets, which blocked a real fix.

Our current model has two routes, in this order.

  1. The official Cloudflare MCP server. Cloudflare publishes official MCP servers, including one that exposes its entire API through a search tool and an execute tool. You authenticate once through the browser with OAuth, so there is no token to store. Register it at user scope and it works in every repository. For background on the protocol, see what is MCP.
  2. One broad-scoped token in the OS keychain. For headless runs, CI, and cron jobs where browser OAuth is unavailable, we keep a single token scoped to all zones on the account. Fetch it from the keychain at run time, never from a plaintext file.

Scope the fallback token to the account's zones broadly, not to one zone, so a new product does not trigger a new token. Include only the permissions you will actually use, for example Pages edit, Workers scripts edit, zone read, DNS edit, firewall edit, zone settings edit, cache purge, and email routing edit.

Why is the command line login not enough?

In our setup, the command line tool's OAuth grant carried read-only zone access. It could deploy a project and could not touch firewall rules or zone settings. That is a permissions fact about our grant at the time, so check what your own login actually carries before relying on it. The lesson holds either way. Find out what each credential can do before an urgent task depends on it.

What was the firewall rule that hid our site from AI answers?

This was the most expensive finding. The site's web application firewall was returning 403 to several AI search crawlers while serving Googlebot a 200. The result was a site that ranked in search but was uncitable by the AI answer engines whose crawlers were blocked. In a browser, everything looked fine.

We found it only after the broad token let us read the zone's rulesets. The fix was a rule change, and the check we run now is mechanical.

  1. List the crawlers you want. Use the current user-agent tokens from each vendor's documentation.
  2. Request a key page as each one. Check the status code, not just the browser view.
  3. Read the firewall rules and bot settings. Look for managed rules and bot toggles that act on automated traffic.
  4. Re-test after every rule change. A rule edit made for spam can block a crawler you want.

Allowing a crawler in robots.txt does not prove it can reach the page. For the full crawler list and what each token does, read what is GEO.

Pro tip

Keep the crawler status check in your release checklist. A firewall change that passes every browser test can still return errors to the bots that decide whether AI assistants can cite you.

Where do AI workloads fit?

Cloudflare's docs also describe Workers AI for running models on its network, and we have used Workers mainly for lightweight glue rather than heavy inference. When a product needs a hosted model, we usually call the model provider from the backend, where we can log, retry, and enforce timeouts consistently. The choice depends on latency needs, model choice, and where your data may go. See choosing an LLM for production for how we compare options.

Frequently asked questions

Should every AI product use Workers?

No. Use a Worker for small server-side logic. Put anything with a database, long-running work, or complex retries on a normal backend.

Is one broad token a security risk?

It is a trade-off. A broad token is powerful, so store it in the OS keychain, never in a repository, and rotate it. Many narrow tokens create their own risk through sprawl and confusion.

How do I check whether AI crawlers can reach my site?

Request a key page using each crawler's user-agent and inspect the status code. A browser test and a robots file check are not enough.

Why use the MCP server instead of the dashboard?

It lets an agent perform and audit changes through one authenticated interface, which is faster and repeatable. You still review each change before it applies.

Want this built for your team? Book a free 30-minute call.

Want to apply this to your business?

Book a free 30-minute call. We will tell you what we would do first.