Skip to main content

11. Models

The Models section allows managing the AI models available in LUCA, including public models from external providers, private custom models, and deployments.

Upon accessing, the left side of the window displays the tabs Public, Private, Deployments, MCPs, and Collections. The Public tab is preselected.

Administration Window Figure 11.1: Administration Window

On the left, the group organizer appears since models, as elements of the application, are also associated with a group.

Depending on the selected tab in the model administration screen, the available models list for each type and a model search field are shown at the top.

The section for Public models allows the management of models that feed from public APIs of major providers.

In the Private tab, the list consists of custom models created from the elements previously deployed in the Deployments tab. This allows for adapting and configuring specific models according to the needs of each group, leveraging private resources.

Public, Private, and Deployments Tabs​

Above the list of models are the following buttons:

Chat

Figure 11.2: Add Button

  • Add: Clicking this button opens a context window that allows choosing the type of element to create: Public, Private, Deployment, MCPs, or Collections. Clicking on one of the first three loads in the central panel the corresponding configuration form (with slight differences in the deployments tab) that allows creating a new model:

Properties Figure 11.3: Properties

  • Properties
    • Model : Select the model's provider using the corresponding icon.
    • Model Type : Text, Image, or Embeddings. Selecting Embeddings simplifies the form: the Prompt and Parameters tabs disappear, as do the MCPs and Collections sections described below (exclusive to Text-type models) and the Include metadata and Group fields described further down.
    • Version : Choose the available version for the model from the dropdown.
    • Name : Assign an identifying name to the model.
    • Endpoint : Indicate the address where calls to the model will be sent.
    • API Key : Enter the key configured for access to the model.
    • Include metadata : By activating this option, the model can access both the data generated during the execution of SQL queries as well as the query itself that was performed. This allows for expanding the available context and improving control over the obtained results, facilitating more precise responses adapted to the information processed in real time. If the option is disabled, the model will only receive the final result of the query, without access to intermediate data or the executed query.
    • Group : Select the group in which the model will be included.
Simplified form for an Embeddings-type model

Figure 11.4: Form for an Embeddings-Type Model

Prompt Figure 11.5: Prompt

  • Prompt
    • System Prompt : Defines the base message that the system will use to contextualize the interaction with the model.
    • User Prompt : Specifies the message that the user inputs as a query or instruction. Default "<|query|>".
    • Assistant Prompt : Indicates the message that the assistant (model) will use as a response or guide. Default "<|context|>".
    • Token Limit : Sets the maximum number of tokens that can be used per week. When disabled, it works without limits.

Parameters Figure 11.6: Parameters

  • Parameters
    • Parameters : Select a parameter from the list.
    • Value : Add a value.
    • Selected Parameters : Displays the previously configured parameters.
note

The available parameters for model configuration are as follows:

  • frequency_penalty: Penalizes the repetition of words or phrases. High values reduce repetitions, low values allow more freedom.
  • logit_bias: Allows favoring or prohibiting certain tokens (words). It is passed as a map of token identifiers and their biases.
  • logprobs: Returns the logarithmic probability of each generated token, useful for debugging and analysis.
  • top_logprobs: Number of alternative tokens and their probabilities to return at each step, along with logprobs.
  • max_completion_tokens: Limits the tokens that the response can generate, controlling the maximum length of the output.
  • n: Number of distinct responses that the model should generate for the same input.
  • presence_penalty: Penalizes the repetition of already mentioned topics, promoting greater variety in the response.
  • response_format: Defines the output format. It can be "text", "json", "xml", etc.
  • stop: List of sequences that, if they appear, automatically stop the generation.
  • temperature: Controls the randomness of the response. Low values (0.0) make the output more deterministic; high values (>1) make it more creative or unpredictable.
  • top_p: An alternative to temperature, limits the options to a subset of tokens whose cumulative probability does not exceed the specified value.

On Text-type models, the MCPs and Collections sections are also added, letting you assign the model the MCP tools and document knowledge bases it can work with.

Assigning MCPs and Collections

Figure 11.7: Assigning MCPs and Collections

Both sections use the same control: a list of available items on the left and one of assigned items on the right, with buttons to move one or several items between the two.

  • MCPs: MCP servers the model can invoke as external tools during the conversation.
  • Collections: Collections of documents the model can use to support its responses.

The Deployment option of the Add button has a particular form.

In the properties section, Model, Version, Name for the model, Server (the client in which the AI model server previously configured in Administration > Clients is deployed), and a Group in which to include it are selected.

Tool Calling: Only shown when the selected Model is the generic option (LUCA icon), used to connect servers not included in the preconfigured provider list. Enables MCP support for the model. Its value depends on the model's family; choose the option matching the deployed model according to each entry's description: hermes (Qwen, DeepSeek, Phi, SmolLM, and most ChatML fine-tunes), llama3_json (Meta Llama 3, 3.1, 3.2, 3.3, 4), mistral (Mistral, Mixtral, Devstral), pythonic (Google Gemma 3+), internlm (InternLM), jamba (AI21 Jamba), xlam (Salesforce xLAM), and granite-20b-fc (IBM Granite).

Deployment Properties Figure 11.8: Deployment Properties

In the parameters, the API Key is entered, and the Hardware to be used for running the model (CPU or GPU) is selected.

Deployment Parameters Figure 11.9: Deployment Parameters

The Deployments tab maintains a simplified add form because model configuration is added in the Connect button, which gives access to a form similar to that of Public and Private. In Restart, the process of stop and start for the respective model is initiated. Once deployed, they appear in the Private tab where they can be marked as default favorite.

Connect and Restart Figure 11.10: Connect and Restart

Chat

Figure 11.11: Edit Button

  • Edit: Modifies the configuration of the highlighted model.
Chat

Figure 11.12: Delete Button

  • Delete: Removes the selected model.
Chat

Figure 11.13: New Chat Button

  • New Chat: Starts a conversation (disabled from the deployments tab) with the selected model from the list.
Chat

Figure 11.14: Reload Button

  • Reload: Updates the list of models.

MCPs​

The MCPs tab lists the configured MCP (Model Context Protocol) servers, which Text-type models can use as external tools.

MCP List

Figure 11.15: MCP List

They are created from the Add > MCPs button:

New MCP

Figure 11.16: New MCP

  • Name and Description: identify the MCP.
  • Group: the group it belongs to.
  • Transport: stdio or http, depending on how the MCP server is accessed.
  • Server Configuration (JSON): the connection configuration, whose format depends on the chosen transport.

For stdio transport:

{
"transport": "stdio",
"command": "",
"args": [],
"env": {}
}

For http transport:

{
"transport": "http",
"url": "",
"headers": {}
}

When editing an existing MCP, a read-only Assigned Models table is also shown, listing the models that have it included in their configuration.

Collections​

The Collections tab lists the document collections that Text-type models can use as a knowledge base to support their responses.

Collection List

Figure 11.17: Collection List

They are created from the Add > Collections button, specifying Name, Description, and Group:

New Collection

Figure 11.18: New Collection

After saving, you access the Documents section, where the files that will make up the collection are uploaded using the Upload Document button.

Documents of a Collection

Figure 11.19: Documents of a Collection

Describe images with AI: Generates a description of each image in the document, using the default favorite of type Image, to include it in the embeddings. Increases upload time.

Knowledge Base​

The Knowledge Base button, located above the Collections list, gives access to a semantic search tool over the content of all of them: typing a question returns the results most similar to the search.

Knowledge base search

Figure 11.20: Knowledge Base Search

Each result shows its source and a text fragment. Document results include a Document button to access the source file; User Memory results (relevant information LUCA has stored from previous conversations) are shown alongside them. Clicking View full fragment opens a window with the full text of the fragment.

Full fragment

Figure 11.21: Full Fragment

The filter icon lets you restrict the search to a specific Collection and adjust the Max. results returned by the search.

Knowledge base filter

Figure 11.22: Results Filter

The Collections button returns to the collections list.

Default Favorite​

On the Public and Private tabs, a model can be marked as the default favorite, so that it is used automatically in certain contexts of LUCA. The selector appears situated above the specifications of each model.

Experts Figure 11.23: Default Favorite

The selector's behavior depends on the Model Type:

  • Text: There can only be one Text-type default favorite. It is the model that responds both in the SQL Console chat (within Queries) and in LUCA's general chat, in the top bar.
  • Image: There can only be one Image-type default favorite, for LUCA features that require a model of this type.
  • Embeddings: In this case the selector is called Global Model, since it is not tied to any specific query or chat type: it is the model LUCA uses to compute the embeddings for Collections.