Generic chat completion in Python transforms

Use GenericChatCompletionRequest with GenericChatCompletionLanguageModel to send text conversations to language models through a common API. The model input selects the model, and the request contains the messages and generation parameters.

This page documents the Python classes from palantir_models and language-model-service-api. For repository setup and model selection, see Use language models within transforms.

Example

Add palantir_models to your repository. Replace <model-rid> with the full resource identifier of a model you can access, and replace the output dataset path. This example uses a lightweight transform.

Copied!
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 import pandas as pd from language_model_service_api.languagemodelservice_api_completion_v3 import ( DialogChatMessage, DialogRole, GenericChatCompletionRequest, ) from palantir_models.models import GenericChatCompletionLanguageModel from palantir_models.transforms import GenericChatCompletionLanguageModelInput from transforms.api import Output, transform @transform.using( model=GenericChatCompletionLanguageModelInput("<model-rid>"), output=Output("/path/to/output/dataset"), ) def compute(model: GenericChatCompletionLanguageModel, output): request = GenericChatCompletionRequest( messages=[ DialogChatMessage( role=DialogRole.SYSTEM, content="Answer questions in one sentence.", ), DialogChatMessage( role=DialogRole.USER, content="Why is the sky blue?", ), ], max_tokens=256, ) response = model.create_chat_completion(request) output.write_table(pd.DataFrame({"completion": [response.completion]}))

GenericChatCompletionLanguageModelInput supplies a GenericChatCompletionLanguageModel to the compute function. The same model input and request types can be used in a Spark transform with @transform(...).

GenericChatCompletionRequest

Import GenericChatCompletionRequest from language_model_service_api.languagemodelservice_api_completion_v3.

Python parameterTypeRequiredDescription
messageslist[DialogChatMessage]YesThe conversation to send to the model, in order. See Message structure.
temperaturefloatNoControls sampling randomness where supported. Supported values and behavior depend on the selected model.
stop_sequenceslist[str]NoSequences that stop generation when encountered, where supported by the model.
max_tokensintNoMaximum number of output tokens to generate for this request. This excludes input tokens and is subject to the selected model's limits.

The optional parameters can be omitted. There is no single set of defaults or supported parameter values across all models; the service and selected model determine behavior for omitted parameters. Start with the required messages and add only parameters supported by your model.

The Python names are stop_sequences and max_tokens. In the serialized request, these fields are named stopSequences and maxTokens.

This request supports text messages and the generation parameters listed above. It does not define image inputs, tool calls, or a structured output schema.

Message structure

Import DialogChatMessage and DialogRole from the same module as GenericChatCompletionRequest.

FieldTypeDescription
roleDialogRoleDialogRole.SYSTEM, DialogRole.USER, or DialogRole.ASSISTANT.
contentstrThe text of the message.

Follow these rules when constructing a conversation:

  • Start with a SYSTEM or USER message.
  • Use SYSTEM only as the first message, if needed.
  • Follow a SYSTEM message with a USER message.
  • Alternate USER and ASSISTANT messages after the optional system message.

For example, a conversation with history can have the roles SYSTEM, USER, ASSISTANT, USER. Include the previous turns in messages when you want the model to use them as context.

GenericChatCompletionResponse

model.create_chat_completion(request) returns a GenericChatCompletionResponse with these fields:

Python attributeTypeDescription
completionstrThe generated text. Read this field to retrieve the answer.
token_usageTokenUsageToken counts and the model's context-window limit.

The token usage object exposes prompt_tokens for input tokens, completion_tokens for output tokens, and max_tokens for the model's context window, which includes both input and output tokens. response.token_usage.max_tokens has a different meaning from the request's max_tokens output limit.

This response has a completion field rather than the choices list used by GptChatCompletionResponse.

Model selection and custom models

Set the model RID on GenericChatCompletionLanguageModelInput. GenericChatCompletionRequest does not have a model or provider field.

For a registered model, use its full ri.language-model-service..registered-model.<id> RID. The Python client uses that identifier to call the registered model. Registering and configuring the model determines which model endpoint handles the request.

When investigating a custom model integration:

  • Confirm that the input uses the intended model RID and that you have access to that model.
  • Confirm that message roles, ordering, and content follow the request structure above.
  • Check the selected model's support for each optional parameter. A common request type does not guarantee identical behavior for every model.
  • Read the generated answer from response.completion.

For guidance on connecting a self-hosted or external model to AIP, see Bring your own model to AIP.