LiteLLM
LiteLLM is a self-hosted AI gateway proxy that provides a unified OpenAI-compatible API for 100+ LLM providers, enabling model fallbacks, load balancing, and spend tracking through a single endpoint.
Categories: Artificial Intelligence
Type: liteLlm/v1
Connections
Version: 1
Bearer Token
Properties
| Name | Label | Type | Description | Required |
|---|---|---|---|---|
| token | API Key | STRING | The master key or virtual key for your LiteLLM proxy. Leave empty if your proxy does not require authentication. | false |
Connection Setup
- Install LiteLLM:
pip install litellm. - Create a config file (
litellm_config.yaml) with your model providers. - Start the proxy:
litellm --config litellm_config.yaml --port 4000. - In ByteChef, set the Base URL to your proxy address (e.g.,
http://localhost:4000/v1). - If your proxy requires authentication, enter your Master Key as the API key.
- Done.
For more details, see the LiteLLM Proxy documentation.
Actions
Ask
Name: ask
Ask anything you want.
Properties
| Name | Label | Type | Description | Required |
|---|---|---|---|---|
| model | Model | STRING Optionsgpt-4o, gpt-4o-mini, gpt-4.1, gpt-4.1-mini, gpt-4.1-nano, o3, o3-mini, o4-mini, claude-sonnet-4-20250514, claude-3-5-sonnet-20241022, claude-3-5-haiku-20241022, gemini-2.5-flash, gemini-2.5-pro, gemini-2.0-flash, deepseek-chat, deepseek-reasoner, mistral-large-latest, codestral-latest | ID of the model to use. The available models depend on your LiteLLM proxy configuration. | true |
| userPrompt | Prompt | STRING | User prompt to the model. | true |
| format | Format | STRING OptionsSIMPLE, ADVANCED | Format of providing the prompt to the model. | true |
| systemPrompt | System Prompt | STRING | System prompt to the model. | false |
| attachments | Attachments | ARRAY Items[FILE_ENTRY] | Only text and image files are supported. Also, only certain models supports images. Please check the documentation. | false |
| messages | Messages | ARRAY Items[{STRING(role), STRING(content), [FILE_ENTRY](attachments)}] | A list of messages comprising the conversation so far. | true |
| response | Response | OBJECT Properties{STRING(responseFormat), STRING(responseSchema)} | The response from the API. | true |
| frequencyPenalty | Frequency Penalty | NUMBER | Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far, decreasing the model's likelihood to repeat the same line verbatim. | false |
| logitBias | Logit Bias | OBJECT Properties{} | Modify the likelihood of specified tokens appearing in the completion. | false |
| logprobs | Logprobs | BOOLEAN Optionstrue, false | Return log probabilities. | false |
| maxCompletionTokens | Max Completion Tokens | INTEGER | Maximum tokens in completion. | false |
| maxTokens | Max Tokens | INTEGER | The maximum number of tokens to generate in the chat completion. | false |
| presencePenalty | Presence Penalty | NUMBER | Number between -2.0 and 2.0. Positive values penalize new tokens based on whether they appear in the text so far, increasing the model's likelihood to talk about new topics. | false |
| reasoning | Reasoning effort | STRING Optionsnone, low, medium, high | Constrains effort on reasoning. Reducing reasoning effort can result in faster responses and fewer tokens used on reasoning in a response. | false |
| seed | Seed | INTEGER | Keeping the same seed would output the same response. | false |
| stop | Stop | ARRAY Items[STRING] | Up to 4 sequences where the API will stop generating further tokens. | false |
| temperature | Temperature | NUMBER | Controls randomness: Higher values will make the output more random, while lower values like will make it more focused and deterministic. | false |
| topLogprobs | Top Logprobs | INTEGER | Number of top log probabilities to return (0-20). | false |
| topK | Top K | INTEGER | Specify the number of token choices the generative uses to generate the next token. | false |
| topP | Top P | NUMBER | An alternative to sampling with temperature, called nucleus sampling, where the model considers the results of the tokens with top_p probability mass. So 0.1 means only the tokens comprising the top 10% probability mass are considered. | false |
| verbosity | Verbosity | STRING Optionslow, medium, high | Adjusts response verbosity. Lower levels yield shorter answers. | false |
| user | User | STRING | A unique identifier representing your end-user, which can help admins to monitor and detect abuse. | false |
Example JSON Structure
{
"label" : "Ask",
"name" : "ask",
"parameters" : {
"model" : "",
"userPrompt" : "",
"format" : "",
"systemPrompt" : "",
"attachments" : [ {
"extension" : "",
"mimeType" : "",
"name" : "",
"url" : ""
} ],
"messages" : [ {
"role" : "",
"content" : "",
"attachments" : [ {
"extension" : "",
"mimeType" : "",
"name" : "",
"url" : ""
} ]
} ],
"response" : {
"responseFormat" : "",
"responseSchema" : ""
},
"frequencyPenalty" : 0.0,
"logitBias" : { },
"logprobs" : false,
"maxCompletionTokens" : 1,
"maxTokens" : 1,
"presencePenalty" : 0.0,
"reasoning" : "",
"seed" : 1,
"stop" : [ "" ],
"temperature" : 0.0,
"topLogprobs" : 1,
"topK" : 1,
"topP" : 0.0,
"verbosity" : "",
"user" : ""
},
"type" : "liteLlm/v1/ask"
}Output
The output for this action is dynamic and may vary depending on the input parameters. To determine the exact structure of the output, you need to execute the action.
How is this guide?
Last updated on