Field Notes · 6 MIN
Why build a content engine in Python instead of using a CMS?
Field notes on a small standard-library Python engine that turns Markdown into SEO-ready pages, with the rules enforced in a rulebook instead of a plugin.
We built one because our needs were narrow, our site was static, and we wanted every page to carry the same structure and metadata without a plugin stack. A content engine here is a single Python script that reads Markdown files and writes finished HTML pages, a blog index, and a sitemap. This post covers what it does, why we kept it deliberately small, and the rules we put in a rulebook instead of in code.
- A small engine that does a few things the same way every time beats a flexible system that drifts page by page.
- Keeping the Markdown subset small is a feature, because it removes formatting decisions that do not belong to the writer.
- Structured data, the table of contents, and the sitemap come from the engine, so no page depends on a human remembering them.
- Quality rules that machines cannot check, such as voice and honesty, live in a rulebook that writers and reviewers read.
- Nactore is an AI-native software engineering partner, and we build internal tools like this when a team needs to move faster.
What does the engine actually do?
You write a Markdown file with a short frontmatter block. You run one command. The engine renders an HTML page in the site's design system and rebuilds the blog index and sitemap. The script uses only Python's standard library, so there is nothing to install and nothing to keep patched.
The generated pages are build artifacts. We never edit them by hand. If a page is wrong, we fix the Markdown and run the engine again. That single rule keeps the source of truth in one place.
| Input | Output |
|---|---|
| One Markdown file per post | One HTML page at a clean URL |
| Frontmatter fields | Title, description, canonical, social tags |
| Headings | An automatic table of contents with anchors |
| Word count | Reading time |
| All posts | Blog index and sitemap, newest first |
Why not use a CMS?
A CMS earns its keep when many non-technical people edit many page types. Our case was the opposite. A small team publishes one kind of page, the site is static, and we care about consistency more than flexibility.
A CMS adds a database, a login surface, plugin updates, and a theme that can change underneath you. Each of those is a place for a page to quietly differ from the last one. A script that reads files has none of that. The trade-off is real. A non-technical editor cannot publish without help, and we accepted that. If your team needs many editors, a CMS may be the right call.
Why keep the Markdown subset small?
The engine supports headings, paragraphs, bold, inline code, links, bullet and numbered lists, tables, blockquotes, and three fenced blocks for a callout, key takeaways, and an FAQ. That is the whole language, on purpose.
A small subset has three benefits.
- Writers stop making layout decisions. They choose structure, and the design system handles presentation.
- The renderer stays small. Fewer features means fewer edge cases and a script a new engineer can read in one sitting.
- Output stays uniform. Every table and callout looks the same on every page.
When a post needs something the subset cannot express, we ask whether the post needs it or whether the idea can be said another way. Usually it is the second.
What does the engine generate that people forget?
Three things that are easy to skip by hand and easy to get wrong.
- Structured data. The engine emits
BlogPostingandBreadcrumbListJSON-LD on every post, andFAQPagewhen an FAQ block exists. Google documents article structured data, and the types are defined at schema.org. Structured data labels what is on the page. It does not create authority. - A table of contents from the headings. Because the contents are generated from the headings, writers are pushed to write headings that read well as a list, and question-shaped headings are easy for answer engines to lift.
- A sitemap that matches the posts. It is rebuilt on every run, so a published post cannot be missing from it.
What goes in the rulebook instead of the code?
Some of the most important rules cannot be checked by a script. We keep them in a written rulebook that every writer reads before a first draft, and a review step checks the output against it.
- Structure. A lead paragraph that answers the title, a takeaways block, five to eight sections, at least one table or list, and a short FAQ.
- Voice. Plain, concrete, written to a peer. Specifics over adjectives.
- Honesty. No invented statistics, client names, or results. External facts link to a primary source.
- A blocklist. Hype words, filler phrases, and self-descriptors we never use, plus a ban on em dashes.
The engine enforces form. The rulebook enforces judgment. Mixing them up is how teams end up with a plugin that checks word counts and a blog that says nothing. For how the same idea applies when AI writes the first draft, see AI evals before production.
When a rule matters and a script can check it cheaply, automate it. When a rule needs judgment, write it down once and put a human review step after it. Do not pretend a regex can judge honesty.
How does this support GEO as well as SEO?
The engine makes pages easier to crawl, parse, and quote. Clean HTML, descriptive headings, a direct answer in the first paragraph, and consistent structured data all help both search engines and AI assistants. None of it guarantees a citation, and we say so in the rulebook. What it does is remove avoidable obstacles and keep every page consistent.
The measurement side is separate. We test what assistants actually say with a harness described in measuring AI answers with Playwright.
What would we do differently?
Two habits we would keep and one change we would make.
- Keep. Standard library only, generated pages never hand-edited, and a small Markdown subset.
- Keep. Publish by pushing to the main branch and letting the host deploy.
- Add. An automated check for banned phrases and dashes before render, since those are mechanical and we currently verify them by search.
Frequently asked questions
Is a custom engine worth it for a small blog?
If the site is static and a technical team publishes it, a small script can be cheaper to own than a CMS. If many non-technical editors publish, a CMS is usually the better fit.
Does structured data improve rankings?
It helps search engines understand the page and may make it eligible for certain features. It does not create authority or guarantee ranking or citation.
Why enforce rules in a rulebook instead of code?
Voice, honesty, and relevance need judgment. Code can check form, such as headings and metadata, but a human review is still needed for substance.
Can an AI agent write posts with this engine?
Yes, if it follows the same rulebook and a human reviews the output. The engine does not care who wrote the Markdown. The rulebook and the review are what keep the quality up.
Want this built for your team? Book a free 30-minute call.
Want to apply this to your business?
Book a free 30-minute call. We will tell you what we would do first.