What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Yes—Databricks AI Functions let you apply a model to rows of data from SQL, so a statement can combine ordinary relational work with tasks such as extracting fields, assigning labels, or searching configured knowledge sources. The best function depends on the task: start with a task-specific function when it fits, and use ai_query when you need more control over the prompt, model, parameters, or output. “One-liner” describes the SQL interface, not the operational work behind it: model latency, compute costs, permissions, licensing, and data governance still matter.
What “one-liner” means in Databricks SQL
Databricks describes AI Functions as built-in functions for applying LLMs and other models to data stored on Databricks. They can be used from Databricks SQL, notebooks, Lakeflow pipelines, and Workflows. The useful idea is to keep filtering, joining, and other relational operations in SQL, then call an AI function where a row needs an AI task.
That can make a complex transformation fit into a SQL statement, but it does not mean the underlying task has become instantaneous or cost-free. The function still invokes model-backed processing, and the workload remains subject to the relevant compute, access, licensing, and governance requirements.
Which AI Function should you use?
Choose by the result you need, not by which function sounds most general. Databricks recommends starting with a task-specific AI Function when one matches the objective; use ai_query when a specialized function does not provide the control or behavior required.
#1 Best Overall
| Function | Best fit and input | Result and control | Status or operating note |
|---|---|---|---|
ai_extract |
Extracting fields from text or parsed-document output, including invoices, contracts, and financial filings. | Structured fields defined by a schema. The schema can describe nested objects and arrays, validate types, and include field descriptions, subject to documented API limits. | Generally available since June 2026. The published default limit is 120 requests per minute per workspace. |
ai_classify |
Assigning text to labels that you supply. | Labels; the API supports label descriptions and multi-label behavior. | Generally available since June 2026. The published default limit is 1,200 requests per minute per workspace. |
ai_parse_document |
Parsing an unstructured document before extraction or other analysis. | Text, tables, figure descriptions, and layout information for downstream processing. | Production status and a throughput limit are not stated in the cited documentation. |
ai_search |
Retrieving information from one or more configured knowledge sources. | Retrieves and deduplicates results, reranks them, and by default synthesizes a grounded answer over the configured sources. | Beta; availability and behavior may change. A throughput limit is not stated in the cited documentation. |
ai_query |
Custom prompts and supported model endpoints when a task-specific function is not enough. | Flexible, prompt-driven output for tasks such as extraction, summarization, or classification, as well as custom ML-serving calls. | Requires Databricks Runtime 15.4 LTS or later; Runtime 18.2 or later is recommended for best performance and the latest features. |
Other documented AI Functions cover sentiment analysis, semantic similarity, summarization, translation, grammar correction, masking, forecasting, anomaly detection, and top-driver analysis. Their inclusion in the catalog does not establish that every function has the same status, runtime requirements, or throughput limit.
When to parse, extract, classify, or search
Parse first when document structure matters
For an unstructured document, ai_parse_document can identify text, tables, figure descriptions, and layout. That parsed output can then feed a later task, such as extraction. Parsing and extraction answer different questions: parsing makes document content and structure available; extraction maps relevant content into fields described by your schema.
Extract when you need fields, not a free-form answer
Use ai_extract when the desired result is a defined record—for example, values represented by a schema rather than an open-ended summary. Schema descriptions and type validation help make the intended shape explicit, while nested objects and arrays support more involved records within the documented API limits.
Classify when the destination is a set of labels
Use ai_classify when each text item should be assigned one or more categories you provide. Label descriptions can clarify the categories, and the API supports multi-label behavior. This is a better match for a bounded labeling task than asking a general prompt to invent a category scheme.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Search when the answer must come from configured knowledge
ai_search retrieves information from one or more knowledge sources. It generates optimized queries, retrieves and deduplicates results, reranks them, and by default synthesizes a grounded answer over those sources. Because the function is Beta, treat its availability and behavior as subject to change; it is not simply a substitute for extraction from a particular row of text.
Use a custom query when you need more control
ai_query is the general-purpose option for a supported model endpoint and a custom prompt. It can handle extraction, summarization, classification, or custom serving calls, but it puts more responsibility on you to define the request and desired output. Databricks recommends choosing a task-specific function first whenever one already matches the job.
Rank #4
What to check before running a batch
- Compute compatibility: AI Functions are not available on Classic SQL warehouses. For
ai_query, Databricks Runtime 15.4 LTS or later is required, with 18.2 or later recommended for best performance and the latest features. Check the specific function’s current documentation for its execution requirements. - Workspace throughput: Databricks’ current AI Functions API reference lists default limits of 1,200 classification requests per minute per workspace and 120 extraction requests per minute per workspace. These are per-workspace request limits, not a guarantee of end-to-end rows processed per minute; batch size and model latency also affect completion time.
- Data access and governance: Confirm that the identity running the SQL can access the input data and the relevant model or knowledge source, and that sending the data for the intended processing complies with your organization’s rules.
- Cost and latency: A concise SQL expression does not remove model execution time or compute costs. Estimate the workload and evaluate it with the actual data volume and chosen function before relying on it in a production pipeline.
- Output validation: For structured extraction, make the schema match downstream expectations and decide how to handle invalid or incomplete values. For custom queries, define and validate the expected response shape rather than assuming the model will always return usable output.
- Availability: Check whether the function is generally available or Beta and verify current behavior before depending on it operationally.
ai_extractandai_classifybecame generally available in June 2026;ai_searchis marked Beta.
Can one SQL statement do the whole workflow?
It can express the relational transformation and an AI operation together when the data, function, and execution environment support that task. A document workflow may still have distinct stages—for example, parsing unstructured content and then extracting schema-defined fields—even if those stages can be composed into a SQL workflow. For a bounded label task, classification is the direct fit; for a grounded answer over knowledge sources, search is the relevant option; for unusual prompt requirements, use a custom query.
In practice, decide the output shape first, then pick the narrowest function that produces it. Confirm the relevant runtime or warehouse support, API limits, permissions, and governance constraints before scaling the statement across a batch.
Quick Recap
Best Value
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

