AI/ML API gives developers shared access to text, image, video and speech models, with a playground, API keys and usage-based billing.
AI/ML API
Explore features, practical uses and pricing below.
AI/ML API gives developers a shared way to call AI models from several providers. Its catalog covers language, image, video, speech, music, embedding and other model families. You choose a model suited to the job, send a request through its documented endpoint and use the response in your own application. The product is useful when a project needs more than one kind of generation and you want to manage access through a common account.
Think of a website that turns a product brief into a description and an illustration. The writing step and the image step have different inputs and return different results. AI/ML API can supply the model access for those steps, while your application still handles the form, saved work, review and final publication. It provides inference infrastructure; the complete product experience remains something you build.
The documentation organizes models by task, developer and supported capabilities. That makes it possible to start with the output you need instead of choosing a model solely because its name is familiar. A text assistant, an image editor and an audio transcription feature should each have their own short list and evaluation criteria.
For a text feature, first decide whether the response needs to be a readable paragraph, a structured record or a conversation. Then inspect the candidate's model reference for the relevant capabilities and parameters. Features such as streaming, structured output, tool calls and file inputs depend on the model. Their presence somewhere in the catalog does not mean that every endpoint accepts the same request.
A receipt-processing feature is a good example. The application needs the merchant, date, currency and total, with an explicit outcome when the image cannot be read. A pleasant explanation of the receipt is a different result. Build a small collection containing clear receipts, blurred photographs, several currencies and a document that is not a receipt. Evaluate whether each candidate produces the fields and failure behavior your interface can use.
For images or video, define the asset before selecting the model. An illustration made from a text prompt is different from editing an uploaded product photograph. A video job also needs an approach to duration, output retrieval and waiting. Consult the specific reference for its inputs and response format; do not reuse a chat-completion request and assume that every media model will understand it.
The same care applies to embeddings. A search feature needs consistent vectors for both documents and queries, along with storage and a retrieval method. Calling an embedding model is one part of that system. Keep the selected model and its vector format consistent when preparing the collection, and plan how existing records would be rebuilt if you later change the model.
AI/ML API's playground is a place to try a model and inspect a request before integrating it. The starter guide describes request examples in Python, JavaScript and cURL. Use this stage to establish one working input, record the settings and inspect the actual result. That gives the developer a known example to reproduce when connecting the backend.
Start with a short, ordinary input from your intended workflow. For a customer-support draft, supply an approved policy excerpt and a sample question. Ask for an answer based on that excerpt and an escalation when it does not contain the answer. Review whether the result follows those instructions before adding a longer conversation or more elaborate prompting.
Once the ordinary case works, add one case the model should decline to resolve from the supplied material. This helps you see whether the proposed feature respects the limits of its information. It also gives the interface designer a concrete empty or uncertain state to accommodate. A successful first response should lead to a repeatable feature, with visible handling for the other outcomes.
Keep a simple record of the model identifier, request parameters, input and result. If the answer changes after a later edit, you can distinguish a changed prompt from a changed model or setting. Saving this information yourself is especially useful during collaboration: another developer can reproduce the same call instead of trying to infer what happened from a screenshot.
The quickstart shows an OpenAI-compatible client configured with the AI/ML API base URL and an AI/ML API key. This is relevant to supported chat-completion integrations, but the model reference remains the authority for the chosen request. Some tasks use other endpoints or generation patterns. Follow their examples rather than treating client compatibility as a promise that every provider feature is identical.
A sensible first integration has a backend endpoint that accepts a narrow, validated input. The backend prepares the model request and returns only the fields the interface needs. Keep the credential on that backend. A browser form should call your application endpoint rather than receive the provider key. This also gives your application a place to enforce limits before a paid model call occurs.
For example, a product-description generator can require a product name, verified attributes and a target audience. Reject an empty brief locally. After generation, show the draft for editing and retain the supplied attributes beside it. If the model adds an unsupported material, certification or delivery promise, the editor can remove it before saving. The API call supplies a draft; the surrounding workflow keeps that draft accountable to the brief.
Handle errors as part of the interface. Distinguish an invalid input from an unavailable service and from a request that takes longer than expected. For an expensive generation, avoid silently submitting the same job several times after a user clicks again. Store the state of the application job and make the retry behavior explicit. These are application decisions that are worth making before a public launch.
AI/ML API documents separate API keys and management keys. Its key controls include descriptive names, endpoint permissions and spending limits with supported reset periods. A management key has a different administrative role from a key used to call models. Keep those roles separate when configuring the project, and consult the key-management guide for the current options.
Name a development key after the project and environment rather than calling every key “test.” When usage increases, that naming helps connect a charge or failed request to the application responsible for it. A separate production key also makes it easier to replace a development credential without interrupting live requests.
Set an initial spending boundary appropriate to the experiment. A feature that is only being evaluated does not need the same allowance as a launched customer workflow. Review what happens when the boundary is reached and make that state understandable to the person using the application. A limit is most useful when it leads to a clear pause or explanation rather than a mysterious blank result.
Keep the full key out of source control, shared notes and client-side bundles. When documenting the integration, record the environment-variable name and purpose rather than the credential. If you need to replace a key, update the running application's configuration and test a harmless request before considering the change complete.
The help center describes AI/ML API billing as pay as you go: an account is funded and model requests consume that balance. Its account guide states that there is no free API plan or recurring subscription fee under that billing structure. Playground access and an API budget are separate considerations. Check the account and billing guide before relying on a free-use assumption.
The price unit depends on the model. Text pricing can involve input and output tokens, while media models can use an image, generation, audio or video-related unit. The pricing interface and model reference help explain which unit applies. Compare the cost of your actual job rather than placing a token price beside a per-generation price and treating them as equivalent.
For a writing feature, record the size of the supplied brief and the expected response. A short label and a lengthy report can have very different usage even when they call the same model. For a media feature, include the parameters that change the charge, such as the output configuration supported by that endpoint. Allow for discarded drafts and retries when estimating a working budget.
A useful trial budget answers a practical question: how much does it cost to obtain one result the user accepts? Run several representative jobs and count the revision attempts. The lowest listed unit price may not produce the lowest accepted-result cost if its outputs repeatedly need to be regenerated. This is a reason to evaluate quality and usage together, using your own examples.
The starter guide describes usage views for spend, requests and tokens, with filters such as key, model and time period. It also explains request logs and CSV exports. These are helpful when you want to connect observed behavior to the requests that produced it, rather than guessing from the account's remaining balance.
If a feature suddenly costs more, begin with a narrow time range and the relevant application key. Check whether the number of requests grew, a model changed or the input and output became longer. A newly added context block can increase the cost of every call without an obvious change in the front-end experience. Compare a recent request with the earlier working example.
When a response fails, inspect the request details and compare them against the model reference. A misspelled model identifier, unsupported parameter or wrong input type requires a different fix from an account-balance problem. Save enough information in your application logs to connect its job to the provider request, while avoiding unnecessary copies of sensitive user content.
For a team, decide who reviews usage and at what interval. The owner of a feature should understand both the model's behavior and its operating cost. A brief review after a release can reveal accidental duplicate submissions or a loop that calls the service more often than intended. Usage visibility is valuable because it helps you correct the application, not simply watch a number rise.
AI/ML API documents batch processing for supported models. A batch has its own submission and status workflow, with asynchronous results rather than an immediate answer for each item. Consult the batch-processing reference for eligible models, request structure, completion behavior and the current pricing conditions.
This approach can suit preparing draft summaries for a collection of public documents overnight. Assign each item an application identifier before submission. When results arrive, associate them with the original records and keep unsuccessful items separate. An editor can then review each draft in context rather than sorting a disconnected list of responses by hand.
Batch work needs a visible job state. Keep track of what was submitted, what is still waiting and which results have been imported. If the surrounding application restarts, it should be able to resume checking an existing batch rather than create an identical one. This bookkeeping is part of your integration; a batch endpoint alone does not establish a finished editorial queue.
For a feature where a user expects an answer while looking at the screen, evaluate the normal request workflow first. A batch may be economical for delayed processing but inconvenient for an interactive form. Choose the execution method around the deadline of the task, and verify the supported models before designing the interface around it.
AI/ML API is worth considering for developers comparing several model providers, teams prototyping features across text and media, and applications that need centralized model access. It is especially relevant when the team is comfortable reading an API reference and building its own validation, review and saved-result workflow.
It is less direct for someone who only wants a ready-made writing editor or finished video workspace. The playground helps with experiments, but a customer-facing product still needs application development. Also evaluate whether a model provider's specialized first-party feature is required: access to a model through an aggregator does not automatically reproduce the provider's entire consumer application.
No. Start from the reference for the selected model and task. Supported chat integrations can use a compatible client, while other generation types have their own inputs, endpoints and result handling.
Yes, its language-model access can supply the generation step. Your application still needs conversation state, input validation, an interface and a decision about when an answer should be reviewed or escalated.
Compare the cost and quality of an accepted result on your own examples. Include the input size, response size, retries and correction work. A low unit price tells only part of that story.
Choose one model and one narrow task. Reproduce a documented example, try representative inputs in the playground, then connect a backend endpoint with a restricted key and a small initial budget.