Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchiTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
To summarize long sales notes in Node.js without sending an oversized request, assemble the full prompt, count its tokens with the selected provider’s request-counting endpoint, and split the text only when it will not fit alongside instructions and a reserved output budget. Summarize meaningful chunks, combine their notes when the account-level context matters, and log actual usage. “Cheap” depends on the model, input and output token rates, request shape, and summary quality; no provider is established as cheapest for an equivalent sales-summary workload.
How do I count tokens before sending a request?
Count the request you intend to send—not just the sales transcript. System instructions, message structure, chunk-specific directions, tools or schemas, and the space needed for the generated summary all affect whether a request fits. Token counts are model- and payload-dependent, so word counts and character counts are rough planning aids, not substitutes for counting the assembled request.
OpenAI documents a Responses input-token endpoint that accepts the same input format as the Responses API and returns an input_tokens count. Its JavaScript SDK exposes the count operation as client.responses.inputTokens.count. The count includes formatting tokens used for request structure. See OpenAI’s token-counting guide.
Count the actual request payload
Build the intended input first, then pass that input to the count endpoint. This avoids estimating the prompt from its visible text while overlooking structure added by the API. If you change the instructions or add chunk metadata afterward, count again: those additions consume tokens too.
A local tokenizer can be useful as a quick plain-text preflight, but OpenAI notes that tools such as tiktoken do not account for images and files, tools, schemas, or every model-specific behavior. OpenAI’s Help Center gives a rough heuristic of about four characters or three-quarters of an English word per token, but tokenization varies with the text and model; use that only for intuition, not a fit decision. The provider’s request-counting documentation is the better guide for the payload you will actually send.
Example: count, then send with the JavaScript SDK
The following pattern uses the documented Responses API workflow. Keep the model name, available context budget, and output reserve in application configuration for the model you select. The numeric budget variables below are illustrative configuration values, not provider limits.
import OpenAI from "openai";
const client = new OpenAI();
const model = process.env.OPENAI_MODEL;
const maxInputTokens = Number(process.env.MAX_INPUT_TOKENS);
const outputReserve = Number(process.env.OUTPUT_RESERVE);
const input = [
{ role: "system", content: "Summarize sales material without inventing facts." },
{ role: "user", content: salesText }
];
const count = await client.responses.inputTokens.count({ model, input });
const availableInput = maxInputTokens - outputReserve;
if (count.input_tokens > availableInput) {
throw new Error("Request needs to be split before summarization.");
}
const response = await client.responses.create({ model, input });
console.log(response.output_text);
console.log(response.usage);
Use the same input shape for counting and generation, and verify the response’s usage fields in the SDK version and endpoint you deploy. The official generation pattern and count example are in OpenAI’s JavaScript text-generation guide and token-counting guide.
Free tools Windows power users keep installed
One-click scans. No signup required.
Rank #2
How do I summarize text that is too long for the model?
Do not assume a long source needs chunking: a model with a large context window may fit it in one request. But when the assembled input plus output reserve does not fit, split the source into semantically coherent pieces, summarize each, and—if a conclusion must reflect the entire account—synthesize the chunk summaries in a final request. Chunking is an engineering strategy, not a guarantee that every qualification or relationship across sections will be retained.
1. Decide whether to split
Compare the count for the full intended request with the selected model’s available context budget after reserving room for the expected output and any model-specific reasoning or output allowance. Context windows are total budgets, not source-text-only limits; OpenAI’s context guide notes that input, output, and potentially reasoning can count toward the window, and that excess tokens may be truncated. See OpenAI’s context and conversation-state guide. Keep the model’s applicable limits and a conservative safety margin in configuration, then validate the choice against actual usage.
2. Split at meaningful boundaries
Prefer paragraph, message, or section boundaries over arbitrary character slices. In sales material, preserve the relationship between a claim and its qualification: for example, keep a promised delivery date with the condition attached to it, or keep an objection with the customer’s explanation. There is no evidence-based universal chunk size in the provider documentation; set chunk size according to the selected model’s budget and the structure of your material.
Rank #3
Attach a source identifier and sequence number to each chunk so the next stage can restore order and trace a summary note back to its source. Avoid cutting an individual message or paragraph just to hit a character target unless that unit itself exceeds the budget.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →3. Summarize chunks for the task
Use a consistent instruction for each piece. If the final output needs an overall account-level conclusion, ask chunk summaries to capture structured facts such as customer needs, objections, commitments, dates, and uncertainty. This makes later synthesis easier to audit than a set of free-form summaries.
4. Combine notes when the whole account matters
Send the ordered chunk notes—not necessarily the original transcript again—to a final synthesis step that distinguishes confirmed facts from uncertain or conflicting details. This stage can identify themes across calls, but chunking may lose cross-section relationships; evaluate the pipeline with representative sales notes and check for omissions, contradictions, and invented commitments. Do not treat automatic truncation as a summarization method: it can remove material without producing a faithful summary.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.How much will this summary API call cost?
For a simple uncached request, estimate cost as (input tokens × current input rate) + (output tokens × current output rate). Apply the rates for the exact model and applicable context tier, and account for cached tokens or other pricing categories only if they apply. Provider pricing is model- and category-specific and can change, so check the live rate page before committing to a model rather than embedding a supposedly evergreen rate in code. OpenAI lists its current rates at the official API pricing page.
A lower input rate alone does not establish a cheaper sales-summary workflow. Chunking can repeat instructions across requests, output lengths vary, and providers may tokenize the same text differently. Measure representative jobs: record request input counts, output usage, and the resulting summary quality. Compare whether the summary retains relevant facts and qualifications, not only the visible response length.
Recommended Free Tools
Compare providers on the request you will send
OpenAI, Google Gemini, and Anthropic each document ways to count request tokens, but the endpoints and count behavior are provider-specific. Use the provider’s own counter on the payload shape you plan to send instead of treating a local count—or another provider’s count—as an exact equivalent.
Best Value
| Provider | Documented counting approach | What to verify |
|---|---|---|
| OpenAI | Responses input-token count endpoint; accepts Responses input format and returns an input-token count. | Confirm the model, request structure, context budget, and current pricing categories. Token-counting guide |
| Google Gemini | Documented count_tokens method and Node.js guidance. |
Check count behavior, limits, and current rates for the specific model and payload. Gemini token guide |
| Anthropic | Documented message token-count endpoint. | Check endpoint constraints, limits, and current rates for the specific model and payload. Token-counting documentation |
Choose based on the specific model’s input and output cost, whether its counter represents the payload you will send, its context and output limits, and the factual quality of summaries on your sales material. Also verify operational requirements such as account limits, latency, data handling, and regional availability directly with the provider; these vary and are not established by token-count documentation. There is no equivalent-workload evidence here that identifies one provider as the cheapest.
What should a production pipeline record?
Log enough to explain both cost and quality without making token estimates stand in for measured usage. A practical record for each job includes:
- Provider, model identifier, and the date or configuration version used for the request.
- Whether the request was sent whole or split, plus chunk count and source sequence identifiers.
- Preflight input count, actual usage returned by the generation API, and output length or token usage when reported.
- Summary checks for missing qualifications, conflicting claims, unsupported commitments, and unresolved uncertainty.
Use observed usage and representative output reviews to tune the safety margin and chunk boundaries. That keeps the pipeline’s cost estimate tied to actual model behavior while preserving the sales details readers of the summary need.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

