Field Notes · 7 MIN
What breaks when you ship an AI product with payments?
Field notes from taking an AI product with live payments to production. Vendor SDKs, silent logging, blocked requests, and why the model is the cheap part.
Almost none of the failures we hit when shipping an AI product with live payments came from the model. They came from the boring parts around it: a payment SDK that would not import, a mail API that rejected requests without explanation, and a web framework that logged nothing in production. This post covers each failure, the fix, and the order we now do things on a new service.
- In our builds, the AI call was the most predictable part of the system. The integration points around it caused the outages.
- A vendor SDK is a dependency with its own failure modes. When the protocol is small, plain HTTP and the standard library are less code than keeping the SDK alive.
- Configure production logging before the first deploy, because a missing traceback makes every other bug harder to find.
- Cost modeling should check the share of each sale the model call consumes before anyone spends time optimizing it.
- Nactore is an AI-native software engineering partner that ships AI features to production with evals, and these notes come from that work.
Why do AI products fail in production for non-AI reasons?
An AI product has the same surface area as any paid web product, plus a model call. Payments, email, auth, and logging all have to work first. The model call is usually one HTTP request with a prompt, and it fails loudly when it fails.
The integrations fail quietly. Every issue below presented the same way: the feature was simply dead in production and perfectly fine locally. That pattern is the signal to stop looking at your prompts and start looking at your deploy.
| Failure | Symptom in production | Fix |
|---|---|---|
| Payment SDK import error | Checkout dead, fine locally | Drop the SDK, use stdlib HTTP and HMAC |
| Mail API returns 403 | Transactional email never arrives | Send a User-Agent header, log the response body |
No logging with DEBUG=False | Exceptions vanish | Explicit LOGGING block in settings |
| Inline comment in env file | Environment flag silently wrong | No comments after values |
Should you use the payment provider's SDK?
Not automatically. In one build, importing the provider's Python SDK raised ModuleNotFoundError: No module named 'pkg_resources' on a modern Python base image. Payments went down completely, and nothing about the code had changed. The SDK depended on a packaging module that newer images no longer ship by default.
The protocol the SDK wrapped turned out to be tiny. Creating an order is one authenticated POST. Verifying a payment is one HMAC. The provider documents the payment signature check as a SHA-256 HMAC over the order ID and payment ID joined by a pipe, keyed with your secret. The webhook validation guide uses the same primitive over the raw request body, and it warns you not to parse or cast the body before verifying.
That maps directly onto the standard library.
``python expected = hmac.new(secret.encode(), f"{order_id}|{payment_id}".encode(), hashlib.sha256).hexdigest() hmac.compare_digest(expected, signature) ``
hmac.compare_digest exists to compare values without leaking timing information, so use it instead of ==. Removing the SDK was strictly less code than keeping it working, and it removed a dependency we did not control.
Why does a mail API return 403 with no explanation?
In the same build, a transactional email API answered every request with HTTP 403 and a bare error code: 1010. The code is a Cloudflare error, and Cloudflare documents 1010 as the site owner banning access based on the browser signature. Our request came from a script using a default HTTP client, with no User-Agent header at all.
Sending any sensible User-Agent made it pass. The second lesson was more useful than the first. We were not logging the response body on HTTPError, so the numeric code alone told us nothing. Log the body of every failed third-party call. It costs one line and saves an afternoon.
Why did production show no errors at all?
Django ships no useful default handler for its own logger once DEBUG=False. The Django logging docs state that the default configuration only displays log records when DEBUG=True. In production, exceptions can disappear without a trace on the console.
This one made the other two hard to find. A dead payment path with no traceback looks like a mystery. The same path with a traceback is a five-minute fix.
Our checklist for a new service, in order:
- Set
LOGGINGfirst. Route thedjangologger and your own loggers to stdout, at a level that includes errors, before the first deploy. - Replace small vendor SDKs. If the protocol is one POST and one HMAC, write it yourself and test it.
- Always send a
User-Agent. Log response bodies on every HTTP error from a third party. - Check env parsing. Our env reader treated the whole line as the value, so
DJANGO_ENV=local # notebroke a flag with no error. Keep comments on their own lines. - Keep secrets in the host, not the repo. Test keys locally, live keys only in the production environment, and let the frontend receive the public key from the create-order response instead of hardcoding it.
Where does the AI cost actually sit?
We ran the numbers on one paid AI product before deciding what to optimize. The model call was a rounding error next to the sale it served, a fraction of one percent of the order value. Acquisition cost and conversion were the real constraints.
That changes where engineering effort goes. Prompt compression and model routing matter at very high volume, or when a single request is expensive. For a product that charges per outcome, they were not the lever. Measure the share of revenue the model consumes first, then decide. For the cases where cost does matter, see LLM cost control and prompt caching.
Before optimizing anything about the model, write down what one successful transaction earns and what the model call costs inside it. If the ratio is tiny, spend the week on the payment path and the funnel instead.
What do we test before taking money?
An AI product that charges money has to survive a hostile first run. We exercise the full path in test mode and confirm each of these before live keys exist.
- Order creation and signature check. Create an order, complete a test payment, and verify the signature server-side.
- Webhook replay. Send the same webhook twice and confirm the second one changes nothing.
- Failure paths. Force a mail failure and a model failure, and confirm the user sees a clear message and the logs show the cause.
- Cold start. Hit the service after an idle period and confirm the first request still completes inside the client timeout.
For how we measure the model side of that path, read AI evals before production.
Frequently asked questions
Should I never use a payment SDK?
That is not what we found. We dropped one SDK because it failed to import on a modern runtime and the underlying protocol was small. If a vendor SDK is maintained, tested against your runtime, and covers a large protocol surface, keeping it can be the right call.
Where should live payment keys be stored?
Only in the production host's environment variables, never in the repository. Local development should use the provider's test keys.
Why verify the webhook against the raw body?
The signature is computed over the exact bytes the provider sent. Parsing and re-serializing the body can change whitespace or key order, and then a valid signature will not match.
Is the model call the expensive part of an AI product?
It can be, but check first. In one paid product it was a tiny share of each sale, and the real costs were elsewhere. Measure before you optimize.
Want this built for your team? Book a free 30-minute call.
Want to apply this to your business?
Book a free 30-minute call. We will tell you what we would do first.