Skip to main content

Frequently Asked Questions (FAQ)

1. What is the difference between B.AI Chat and the API?​

Chat lets you converse directly through the website or app. The API lets developers integrate models into their applications or tools. Available models, features, and billing rules may differ; see the relevant documentation for each service.

2. How do I obtain and use an API Key?​

Create an API Key in the B.AI console, then configure it in your application or client. Authenticate requests using Authorization: Bearer <API Key> or x-api-key. Keep your key private and avoid exposing it in screenshots, logs, or source repositories. See the API Reference.

3. Which API Base URL should I use?​

The production Base URL is https://api.b.ai/v1. Choose the request endpoint according to your client's protocol; supported endpoints may vary by model. See the API Reference.

4. What can I do if I cannot access B.AI from mainland China?​

Open B.AI Service Availability in the B.AI Android app to find currently available web addresses and API Base URLs. Copy the displayed address in full. See App and Service Availability.

5. Where can I download the B.AI app?​

Install it through Google Play, or scan the Android download QR code in the upper-right corner of the B.AI website to download the APK. Android 8.0 or later is supported. See App and Service Availability.

6. What if I do not know which model to choose?​

Select Auto in Chat to let B.AI choose an available model for each task. You can select a model manually when you need a fixed model or want to test its capabilities. Auto has no additional selection fee; billing follows the model actually used and its usage. See Auto Mode.

7. How do I fill in the model name in an API request?​

Open the API Key page, find the model you want to use, and select the copy button to the right of its name. Paste the copied model ID into the request's model field.

You can also use GET /v1/models to retrieve the models available to your current API Key. Confirm that the selected model supports the endpoint you are using. See the API Reference.

8. How are input, output, cache writes, and cache reads billed?​

Input is the content sent to the model, and output is the content it generates. Cache writes create reusable cached content, while cache reads reuse it in later requests. Rates, cache durations, and billing behavior vary by model. Some models also count reasoning tokens as output usage. See the relevant model details.

9. How do explicit and implicit caching differ? Are there extra charges?​

Explicit caching typically involves caches created or specified by the caller. Implicit caching automatically identifies reusable content. Both may incur cache write or read charges, and some models also charge for explicit cache storage based on its duration. See the model details for supported caching options and rates.

10. Does web search incur an additional charge?​

Web search is billed separately per use, with rates varying by model. Search results included in the model's context may also incur token charges. Check the model's web search pricing and your usage records.

11. Why might my actual charge differ from the documented price?​

The documentation shows standard reference prices. Actual settlement may also depend on promotional discounts, account benefits, cache hits, context tiers, time-based pricing, and tool usage. See Pricing and Usage and Promotions and Pricing Updates.

12. How do long-context rates and DeepSeek Peak and Off-Peak pricing work?​

Some models apply different rates when input exceeds a specified length. See the model details for the threshold and which parts of the request those rates cover.

For DeepSeek models with time-based pricing, Peak periods are Monday through Friday, excluding Chinese statutory public holidays, from 09:00-12:00 and 14:00-18:00 Beijing Time (UTC+8). All other periods, including weekends and Chinese statutory public holidays, are Off-Peak. Off-Peak prices are half the Peak prices. B.AI Chat always uses Off-Peak pricing for these models. See Pricing and Usage.

13. How do I check my balance and actual charges?​

Use the platform's Usage page to view your balance, the model actually used, token usage, and Credits consumed. Developers can also use the Balance API to query the balance or quota available to the current API Key.

14. Why were Credits deducted when the model did not provide an answer?​

Model charges are based on actual billable usage, including input, output, cache writes, and cache reads. Having no visible answer does not necessarily mean there was no usage: the request may already have processed input tokens, and some reasoning models may generate reasoning tokens that are not displayed to the user. Such usage may still be billable.

Check the usage records on the Usage page for the actual charge. If you have a question about a charge, keep the request time and relevant records to help investigate it.

15. How should I troubleshoot API errors?​

Read the error message in the response, then check the following based on its status code:

Status codeWhat to check
400Parameters, request format, and model compatibility with the endpoint
401API Key and authentication headers
403Account status, subscription, or model permissions
404Request path and model ID
429Request rate and concurrency; reduce concurrency and retry with exponential backoff
500 / 502 / 503Retry later and retain the request ID for investigation

See the API Reference.

16. Why can a request fail even when the input is below the model's stated context limit?​

B.AI does not impose an additional context-length limit. Requests must still follow the rules of the model and its upstream API.

The stated context capacity is not necessarily available entirely for the main input text. System prompts, conversation history, and tool-related content also consume space, and some APIs reserve capacity for output. Client-side token estimates may also differ from server-side counts, so the length of the main text alone cannot determine whether a request exceeds the limit.

Trim conversation history, set the maximum output length according to your needs, and leave some context capacity available. If length-related errors persist, split oversized content into smaller parts or extract the key information before submitting it.

17. Why does the Responses API report cache read usage but no cache write usage?​

Some models use implicit caching, which is managed automatically by the upstream service without explicit client configuration. Usage fields may also vary between upstream services, and some responses include only cache read usage without a separate cache write field.

A missing write field does not mean caching was ineffective. Cache read usage greater than 0 indicates that the request hit the cache, but the response alone cannot determine how much new content was written to the cache.