Credits
Credits are the billing unit for using AI features. Whenever content is processed, generated, transcribed, analysed, or structured, credits are deducted.
How many credits are consumed depends on:
- which model is in use
- how much is processed
- which type of processing takes place – e.g. text processing, reasoning, OCR, embedding, or audio processing
The current credit values and model assignments can be found in the central credit overview in the Google Sheet.
Overview
| Modality | Typical billing |
|---|---|
| Text | by input, output, and optionally reasoning effort |
| Audio | per voice input or per minute (Konversa) |
| Images | flat rate per action (Image-to-Text) or by quality (image generation) |
| Documents | depending on text processing and OCR usage |
| Embeddings | by input tokens |
How the system works
Credits are not billed as a flat rate per product, but according to the type of processing.
Depending on the use case, different billing logics may apply:
- Text – when processing and generating text
- Reasoning Effort – when a model uses additional computational effort for in-depth reasoning
- Audio – for voice inputs, transcriptions, or meeting processing
- Images – for image generation or extracting content from images
- OCR – when scanned documents or images first need to be made machine-readable
- Embeddings – when text is converted into vectors, e.g. for search or knowledge access
Text
For text-based AI features, credits are calculated based on the amount of text processed and the computational logic used.
Three components are particularly relevant:
- Input – the text passed to the model
- Output – the text the model returns
- Reasoning Effort – additional computational effort for models that use in-depth reasoning
Not every model uses reasoning. When it is used, it can additionally increase credit consumption.
Examples
| Use case | Billed components |
|---|---|
| Text request in chat – ask a question, receive an answer | Input + Output |
| Text request with reasoning model | Input + Output + Reasoning Effort |
| Summarise a document | Input (document) + Output (summary) |
Audio
In the audio area, a distinction is made between dictation in chat and Konversa.
Dictation in chat
When speech is converted directly into text in the chat, billing is currently per voice input.
Konversa
Konversa is used for audio and meeting processing. Billing depends on how the recording is created and processed.
Online meetings (Microsoft Teams, Google Meet, Zoom)
For online meetings with a meeting bot, two separate credit positions apply:
- Bot minutes – credits per minute that the note taker actively participates in the meeting
- Transcription minutes – separate credits for transcription, depending on the selected data processing system
File upload (voice file)
When an audio file is uploaded directly, bot minutes do not apply. Only transcription minutes are billed, since no online meeting takes place.
In-person meetings
For recordings created directly via the browser, the same billing applies as for file upload: only transcription minutes.
Additional text processing
When further content is generated from a transcription (e.g. summaries), additional text costs apply.
Examples
| Use case | Billed components |
|---|---|
| Online meeting with bot | Bot minutes + Transcription minutes (+ text processing) |
| Upload an audio file | Transcription minutes (+ text processing) |
| In-person meeting via browser | Transcription minutes (+ text processing) |
→ Learn more: Dictations, In-Person Meetings, Online Meetings
Images
Image generation
For image generation, credit consumption depends on two factors:
- Quality – LOW, MEDIUM, or HIGH
- Format – square, portrait, or landscape
The general rule is: higher quality and larger formats consume more credits.
| Quality | Format | Resolution | Credits |
|---|---|---|---|
| LOW | Square | 1024×1024 | 1.0115 |
| LOW | Portrait | 1024×1536 | 1.4713 |
| LOW | Landscape | 1536×1024 | 1.4713 |
| MEDIUM | Square | 1024×1024 | 3.8621 |
| MEDIUM | Portrait | 1024×1536 | 5.7931 |
| MEDIUM | Landscape | 1536×1024 | 5.7931 |
| HIGH | Square | 1024×1024 | 15.3563 |
| HIGH | Portrait | 1024×1536 | 22.9885 |
| HIGH | Landscape | 1536×1024 | 22.9885 |
The values shown are theoretical reference values.
Image-to-Text
When information is extracted from images, credit consumption is billed as a flat rate per action.
Examples
| Use case | Billed components |
|---|---|
| Generate an image (e.g. in chat) | Text input (prompt) + image generation by quality |
| Upload an image and extract content | Flat rate per action |
Documents
For documents, the decisive factor is which type of processing actually takes place.
- If a document already contains a readable text layer, the normal text costs apply for the processed text.
- If content first needs to be extracted from a scanned document or image, OCR is additionally used.
Depending on the use case, text costs and OCR costs may therefore be combined.
Examples
| Use case | Billed components |
|---|---|
| Process a document with text layer | Text costs (Input + Output) |
| Process a scanned document | OCR + Text costs |
Embeddings
Embeddings are used to convert text into vectors – for example, for search, knowledge access, or semantic processing.
The input text is processed. Billing is therefore based on input tokens.
Example
| Use case | Billed components |
|---|---|
| Vectorise text for search or knowledge access | Input tokens |
Further details
The complete and current overview of modalities, models, and credit values can be found in the Google Sheet with the credit overview.
You can view your organisation's credit consumption in the Analytics Dashboard.