QUICK ANSWER
File Token Counter
Count the processed model input, not the file’s byte size. An AI provider may extract text, render PDF pages as images, summarize spreadsheet structure, or ignore unsupported embedded content. The relevant token total comes from sending the actual supported file through the counting method for the same provider and model you plan to use.
Why file size is not a token count
File bytes include containers, compression, formatting, fonts, metadata, and media. Tokens describe the input representation processed by a model. Two files with similar byte sizes can therefore produce very different token totals, while the same file can be processed differently by different APIs.
How common file types are processed
OpenAI documents different processing paths for file inputs. PDF requests can include extracted text and page images. Text, code, rich-document, and presentation files are processed primarily through extracted text. Spreadsheet inputs use a spreadsheet-specific flow rather than sending every cell as an unchanged text stream.
How to get the count you need
Choose the provider and model first. Then submit the same supported file, prompt, tools, and processing options to that provider’s token-counting endpoint. For OpenAI Responses API inputs, the input-token endpoint accepts supported files and reports the model’s processed input count before generation.
When a text tokenizer is enough
If your application extracts plain text locally and sends only that text to the model, count the exact extracted string with the matching tokenizer. The OpenAI Token Counter can measure OpenAI-encoded text, but its result does not include file parsing, images, request wrappers, or other structured content. Use the Token Cost Kit calculator after you have the relevant input and output totals.
Official reference
OpenAI’s file-input guide lists supported formats and explains how file types are processed. Its token-counting guide describes counting supported file inputs before sending a response request.