Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Set the Gemini API’s thinking_level to low, medium, or high to control how much reasoning Gemini 3.8 Flash applies. Google documents medium as the default. Start there for general workloads, use low when speed and token spend matter most, and reserve high for tasks where deeper, multi-step reasoning is worth potentially longer waits and higher token use.

These levels are qualitative controls, not promises of a particular response time, output length, or accuracy. Google’s documentation does not publish latency benchmarks by level, so choose a production setting by measuring your own TypeScript application.

Set the thinking level in a TypeScript project

Google’s JavaScript example for the Gemini Interactions API uses the @google/genai SDK. The same JavaScript-compatible request shape can be used from TypeScript:

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import { GoogleGenAI } from "@google/genai";

const client = new GoogleGenAI({});

const interaction = await client.interactions.create({
  model: "gemini-3.8-flash",
  input: "Summarize this incident report and identify its unresolved causes.",
  generation_config: {
    thinking_level: "low",
  },
});

console.log(interaction.output_text);

This follows Google’s JavaScript documentation. The documented model ID is gemini-3.8-flash, and Google lists the model as stable. The available levels are low, medium, and high; minimal is not supported and returns an error. Check the installed SDK’s current typings and API availability: the documentation establishes the request shape, but not a version-specific TypeScript declaration or compiler requirement. See Google’s Gemini 3.8 Flash model reference.

Choose a level for the work, not a promised speedup

Google describes Gemini thinking as dynamic: the model adjusts reasoning to the complexity of a request. The setting influences reasoning depth, but it does not guarantee a fixed time-to-answer, token count, or quality level.

Level Best fit Trade-off
low Latency-sensitive, routine tasks such as real-time chat, draft writing, and fast data analysis Google says it reduces time-to-answer for these use cases; it may be a poor fit when a task needs deeper reasoning.
medium General use; Google describes it as the balance suited to most tasks and recommends it for complex coding and agentic use cases Documented default; a practical baseline before workload-specific tuning.
high Difficult multi-step reasoning, mathematics, or tasks where deeper reasoning and tool orchestration matter May involve longer waits and more token use.

Google’s qualitative guidance appears in its Gemini 3.8 Flash model guide. For a consequential task, consider error cost and tool-call reliability alongside latency and billing: a fast response is not useful if it fails the task, while deeper reasoning may not justify its added cost for routine requests.

Rank #2
TypeScript Programming Language - Software Engineer & Coder T-Shirt
  • TypeScript implements a superset of syntax for strictly typed development, facilitating deep static analysis and enhanced development environment integration. The compiler translates source into standard script formats, ensuring parity across any runtime.
  • TypeScript is ideal for front-end developers, full-stack engineers, and software architects who build large-scale web applications. It serves those looking to improve code excellence, reduce bugs through static checking, and maintain complex projects more.
  • Lightweight, Classic fit, Double-needle sleeve and bottom hem

Benchmark your application before setting a production default

Because Google does not provide comparable numeric latency or accuracy results for each level, run a controlled comparison on prompts representative of your users’ work.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. Prepare a representative prompt set. Include routine requests, harder multi-step requests, and cases that use tools if your application uses them.
  2. Run each prompt at each level. Keep the model, input, and other request settings consistent; change only thinking_level.
  3. Measure the whole outcome. Record end-to-end latency, billed input and output tokens, task success, and errors or tool-call failures. Assess response completeness as well as speed.
  4. Select by workload. Use the level that meets your quality and reliability needs within your latency and cost constraints. Different request types can use different levels.

This process is application-specific advice, not a claim that any level has a measured speedup or accuracy advantage. Re-run it when prompts, SDKs, models, or workload patterns change.

Reduce reasoning effort instead of using a tiny output cap

Google recommends lowering thinking_level to low or medium to reduce cost or latency without truncating responses. A small max_output_tokens limit is not an equivalent control: it is a hard cap that includes thought tokens and can stop generation while the model is reasoning, leaving an incomplete or empty answer while still billing for generated thinking tokens. See Google’s thinking guide.

Understand how thinking affects API cost

Google’s pricing page says output pricing includes thinking tokens, so internal reasoning can affect the bill even when the final answer is short. Published standard paid-tier rates for Gemini 3.8 Flash are scheduled to change on January 1, 2027:

Period Input per 1 million tokens Output per 1 million tokens
Through December 31, 2026 $0.75 $3.75
Starting January 1, 2027 $1.50 $7.50

These are Google’s published standard paid-tier rates, not a cost estimate for a particular prompt; actual spend depends on tokens consumed and service tier. Google also lists Batch and Flex at half the standard rates during the introductory period, subject to their terms. Batch is intended for asynchronous processing; Flex offers lower prices with variable latency and best-effort availability. Check the current Gemini Developer API pricing before deployment because prices and terms can change.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Account for the model’s token limits

Google lists a maximum of 1,048,576 input tokens and 65,536 output tokens for Gemini 3.8 Flash. The output limit includes thinking tokens, so the configured cap can constrain reasoning as well as the visible answer. These are model limits, not recommended targets for ordinary requests; choose an output cap that leaves room for the response you need, and use thinking_level to adjust reasoning effort.

When the compatibility interface is relevant

If an application uses Google’s OpenAI-compatible interface rather than the native Gemini SDK, Google documents a mapping from reasoning_effort to Gemini’s thinking_level. For a TypeScript project using the native @google/genai example above, set generation_config.thinking_level directly. See Google’s OpenAI compatibility guide.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.