Field Notes · 6 MIN

What LLM settings matter for temperature and cost?

Measured notes on temperature, hidden reasoning tokens, and JSON mode for a long-form generation task. Sweep your own settings, the docs default may be wrong for you.

By NactorePublished 9 Aug 2026All articles

Three settings moved quality and cost more than anything else in one long-form generation feature we built. A lower temperature than the documented suggestion, reasoning turned off, and a deliberate decision about JSON mode. None of them were discoverable without measuring. This post gives the findings, the method, and the caveat that matters most, which is that your task may land somewhere different.

Key takeaways
  • A documented default temperature is a starting point. In our sweep, values above 1.0 degraded into incoherent text after the first paragraph.
  • Reasoning mode that is on by default can bill as output tokens, so leaving it on for a task that does not need it quietly raises cost.
  • JSON mode shortened our output by roughly a quarter. That is a penalty for long prose and a feature for short prose.
  • Never trust the model's echo of data you already hold. Use your own source of truth for names and identifiers.
  • Nactore is an AI-native software engineering partner, and we tune model settings against measured output, not defaults.

What did we measure?

We called a hosted model through its OpenAI-compatible chat completions endpoint, using plain urllib and no SDK. The task was generating a structured, readable piece of writing from a fixed input. We changed one setting at a time and read the outputs.

The settings below are what held up. They are measurements on one task with one model, not general truths.

SettingWhat we usedWhy
Temperature0.7Higher values degraded after the first paragraph
ReasoningDisabledNot needed for the task, and it adds output cost
JSON modeEnabledShortens output, which suited this task
TransportPlain urllibNo SDK to break on a runtime upgrade

Is the documented temperature the right one?

Not necessarily. The documentation suggested a higher temperature for creative writing. We swept four values, 0.7, 1.0, 1.3, and 1.5, and read the full outputs. Everything above 1.0 collapsed into incoherent text after the first paragraph. The first paragraph looked fine at every setting, which is exactly why a quick spot check would have missed it.

That points to the method more than the number.

  1. Fix the input. Use the same prompt and data for every run.
  2. Sweep the setting. Change only temperature, and run several samples per value.
  3. Read the whole output. Failures can appear late, so reading only the opening misleads you.
  4. Pick the highest value that stays coherent. Higher can mean more varied, so take it only while quality holds.
  5. Re-measure per task. A different task, such as short answers or code, can land elsewhere.

The takeaway is not "use 0.7." It is "do not trust a default you did not test." For a broader approach to quality measurement, read how to evaluate LLM output quality.

Why can reasoning mode raise your bill?

Some current models have a thinking or reasoning mode on by default. The DeepSeek thinking mode documentation states that thinking is enabled by default and explains how to turn it off with a thinking parameter set to disabled. The same page notes that temperature has no effect in thinking mode, so a temperature sweep means little until you know which mode you are in.

In our case, reasoning tokens bill as output tokens. We saw this in the provider's billing behavior for our usage, so confirm it on your account before you rely on it. Our task did not need step-by-step reasoning, so leaving it on would have multiplied cost for no quality gain.

The rule we took from this is simple. For every model call, decide explicitly whether the task needs reasoning. Turn it off where it does not, and measure where it does. For more on controlling spend, see LLM cost control and prompt caching.

What does JSON mode do to output length?

Requesting a JSON object response cost us about 26 percent of output length in our measurements. Whether that is bad depends on the task.

  • Long prose. It is a penalty. If you want a long, flowing answer, JSON mode can shorten and flatten it.
  • Short structured output. It is a feature. If you need a compact, parseable result, the shorter output is exactly what you want.

We ended up with it enabled because our output was meant to be short and structured. Decide by the shape of the output you want, not by habit. For parsing reliability in general, see structured output and JSON reliability.

Should you trust the model's echo of your data?

No. We told the model which items were in play, and it still sometimes renamed them in its output. So the names shown to users come from our own data, not from the model's response.

Validate the shape of every response yourself and treat the model as a writer, not a source of record. Anything with a canonical value, such as a name, an ID, or a price, should be joined from your own store after generation.

Pro tip

Make a one-page settings sheet for every model call in your product. List the model, temperature, reasoning mode, output format, and the measurement behind each choice. When a vendor changes a default, you will know what to re-test.

Why use plain HTTP instead of an SDK?

For one chat completions call, a plain HTTP request is enough. It removes a dependency, makes the request visible, and avoids breakage when a runtime upgrades. This is the same trade we describe in what breaks when you ship an AI product with payments. If you need streaming, tool calling, and retries across several providers, an SDK or a gateway may earn its place.

What should you take away?

Defaults are a vendor's guess about an average task. Your task is not average. Measure temperature per task, decide on reasoning explicitly, choose JSON mode by output shape, and keep your own data as the source of record. Then write down what you measured, and re-check when the model or its defaults change.

Frequently asked questions

Is a lower temperature always better?

No. It depends on the task. We took the highest temperature that stayed coherent for our use, and that was lower than the documented suggestion. Measure your own.

How do I know if reasoning tokens are being billed?

Check your provider's usage fields and billing documentation for your account, then compare usage with reasoning on and off for the same prompts.

Does JSON mode hurt quality?

It can shorten output. For short structured results that is helpful, and for long prose it may be a drawback, so test with your own prompts.

Why not trust the model to repeat names correctly?

Models sometimes rename or alter items they were given. Join canonical values from your own data after generation.

Want this built for your team? Book a free 30-minute call.

Want to apply this to your business?

Book a free 30-minute call. We will tell you what we would do first.