QUICK ANSWER
Gemini Token Counter
Use the Gemini API’s countTokens method to count input before generation. Send the same Gemini model and content you plan to use in the real request. The method runs that model’s tokenizer and returns the input total; after generation, the response usage metadata provides the actual input and output usage.
Why the Gemini model must be specified
The token-counting request identifies a Gemini model because tokenization belongs to the model. A count should therefore be made with the same model selected for generation rather than borrowed from another provider or model family.
What Gemini can count
Gemini tokenization covers more than plain text. Google documents token counting for chat content, system instructions, tools, cached content, images, audio, video, and PDFs. To measure a multimodal request, include the same supported media and instructions that the generation request will receive.
Input count versus final usage
countTokens returns the input count before a response exists. It cannot predict how long Gemini’s answer will be. After generation, inspect the returned usage metadata for the recorded input, output, cached, thought, tool-use, and total fields that apply to that response.
Use the right result
Use Google’s official Gemini token guide for request counting. Once you have the relevant totals, use the Token Cost Kit calculator for token-based cost planning. For OpenAI text encodings, use the separate OpenAI Token Counter.