Databricks

Connect Foundry to Databricks to leverage a range of capabilities on top of data, compute, and models available within Databricks.

Supported capabilities

CapabilityStatus
Exploration🟢 Generally available
Bulk import🟢 Generally available
Incremental🟢 Generally available
Virtual tables🟢 Generally available
Compute pushdown🟢 Generally available: Python transforms, Pipeline Builder
External models🟢 Generally available

When external access is enabled in Unity Catalog, the Databricks connector exposes additional Delta Lake and Apache Iceberg functionality through virtual tables. Refer to the Virtual tables section for configuration details.

Setup

  1. Open the Data Connection application and select + New Source in the upper right corner of the screen.
  2. Select Databricks from the available connector types.
  3. Follow the additional configuration prompts to continue the setup of your connector using the information in the sections below.

Learn more about setting up a connector in Foundry.

Connection details

The following configuration options are available for the Databricks connector:

OptionRequired?DefaultDescription
HostnameYesThe hostname of the Databricks workspace.
HTTP PathYesThe Databricks compute resource’s HTTP Path value. This can be either a:
  • SQL warehouse (recommended) of form /sql/<version>/warehouses/<warehouseId>
  • Compute cluster of form sql/protocolv1/o/<workspaceId>/<clusterId>
Cloud FetchYesTrueIndicates whether Cloud Fetch should be enabled. Refer to the Networking documentation below to ensure suitable network connectivity to the cloud storage locations.

Refer to the official Databricks documentation ↗ for information on how to obtain these values.

Authentication

You can authenticate with Databricks in the following ways:

MethodDescriptionDocumentation
Basic authentication [Legacy]Authenticate with a user account using username and password. Basic authentication is legacy and not recommended in production.Basic authentication ↗
OAuth machine-to-machineAuthenticate as a service principal using OAuth. Create a service principal in Databricks and generate an OAuth secret to obtain a client ID and secret.OAuth for service principals (OAuth M2M) ↗
Personal access tokenAuthenticate as a user or service principal using a personal access token.Personal access tokens (PAT) ↗.
Workload identity federation [Recommended]Authenticate as a service principal using workload identity federation. Workload identity federation allows workloads running in Foundry to access Databricks APIs without the need for Databricks secrets. Create a service principal federation policy in Databricks and follow the displayed instructions to allow the source to securely authenticate as a service principal.Databricks OAuth token federation ↗

Refer to our OIDC documentation for an overview of how OpenID Connect (OIDC) is supported in Foundry.

Refer to Federation policy requirements for the exact policy values expected by Foundry.

For full feature support, ensure that the credentials provided have been granted the relevant privileges on the relevant catalog and compute resources.

When using OAuth machine-to-machine authentication in Azure Databricks, be sure to use your Databricks service principal client ID and secret. Authentication using a Microsoft Entra service principal is not supported. Learn more about service principals in Azure Databricks. ↗

Federation policy requirements

When using Workload identity federation, Foundry issues an OIDC identity token that Databricks trusts through a federation policy. Create the policy in Databricks with the following values, which are also displayed in the source configuration panel when you select this authentication method:

Policy fieldValue
IssuerThe Foundry OIDC issuer URL. Foundry publishes its signing keys and metadata at <issuer>/.well-known/openid-configuration.
AudienceThe audience configured on the Foundry source.
SubjectThe RID of the Foundry source, in the form ri.magritte..source.<id>.
Subject claimLeave as the default value of sub.
Token signature validationLeave unset so that Databricks discovers the signing keys from the Foundry well-known endpoint.

You must create a service principal federation policy rather than an account-wide federation policy. Account-wide policies expect the sub claim to contain a Databricks username, but Foundry sets sub to the source RID, so an account-wide policy will never match.

Refer to the official Databricks documentation ↗ for details on how to create a federation policy.

These values are also available in code as the issuer_url, audience, and subject properties of source_configuration.authentication, alongside the service_principal_application_id of the service principal the policy is attached to. Refer to Use Databricks sources in code for an example.

Networking

For Databricks connections, add the appropriate egress policies when setting up the source in the Data Connection application.

Databricks connections typically open a large number of connections at the same time. When using agent proxy egress policies, you may exhaust the connection pool on your agent. If you experience connection pool errors, increase the maxConnections and coreConnections settings in your agent proxy configuration.

The Databricks connector requires network access to the Hostname provided in the configuration options on port 443. This grants access for Foundry to connect to the Databricks workspace and Unity Catalog REST APIs.

Cloud Fetch

Cloud Fetch is a feature of the Databricks JDBC driver. Cloud Fetch enables parallel data extraction from Databricks to Foundry through cloud storage, delivering up to 10x faster performance compared to traditional single-threaded transfers.

When enabled, additional network policies may be needed to allow outbound connections to the cloud storage service (AWS S3, Azure Data Lake Storage, or Google Cloud Storage) where Databricks temporarily stores query results. If you are using a Foundry worker connection, egress policies will need to be created for the workspace storage bucket. This is the cloud storage location used by Cloud Fetch.

Review Databricks' official documentation ↗ for details.

External access to storage locations (virtual tables only)

The Virtual Tables section of this documentation provides details on external access in Unity Catalog and the functionality it enables. External access requires network connectivity to a table's storage location (managed or external). Egress policies will need to be created for each storage location to benefit from the features enabled by external access.

Egress policies only cover traffic leaving Foundry; you should also ensure that any network controls on the storage location permit traffic from Foundry. These controls vary depending on the cloud provider. Learn more about identifying the IPs where Foundry traffic originates.

Refer to the official Databricks documentation ↗ for more information on external access and how to determine the storage locations of tables.

Examples

Below we provide example egress policies that may need to be configured to ensure network connectivity to Databricks.

TypeURLDNSPort
Databricks workspacehttps://adb-5555555555555555.19.azuredatabricks.net/adb-5555555555555555.19.azuredatabricks.net443
Azure storage location [1]abfss://<container-name>@<account-name>.dfs.core.windows.net/<table-directory><account-name>.dfs.core.windows.net

<account-name>.blob.core.windows.net
443
Google Cloud Storage (GCS) storage locationgs://<bucket-path>/<table-directory>storage.googleapis.com443
S3 storage locations3://<bucket-path>/<table-directory><bucket-path>.s3.<region>.amazonaws.com443

[1] Be sure to include both blob.core.<endpoint> and dfs.core.<endpoint> domains when configuring access to Azure storage locations. endpoint may vary depending on the Azure Cloud environment.

For egress policies that depend on an S3 bucket in the same region as your Foundry instance, ensure you have completed the additional configuration steps detailed in our Amazon S3 bucket policy documentation for the affected bucket(s).

In a limited number of cases, depending on your Foundry and Databricks environments, it may be necessary to establish a connection via PrivateLink. This is typically the case where both Foundry and Databricks are hosted by the same CSP (for example, AWS-AWS or Azure-Azure). If you believe this applies to your setup, contact your Palantir representative for additional guidance.

More options: SSL and hostname validation

You may additionally need to pass in a JDBC property to allow self-signed certificates.

How to identify if this property is needed:

  • SSL connections validate server certificates. Normally, SSL validations happen through a certificate chain. By default, both agent and Foundry workers trust most industry-standard certificate chains.
  • The server must provide the full certificate chain in order for SSL verification to work. To obtain the certificate chain for the Databricks server, run the command openssl s_client -connect {hostname}:{port} -showcerts, then verify the chain using the OpenSSL command line utility or any other available tool.
  • If the server to which you are connecting has a self-signed certificate, or if a firewall performs TLS interception on the connection, the connector must trust the certificate. Learn more about using certificates in agent-based connections.
  • If you are creating a Foundry worker connection and are using a self-signed certificate, you will need to add a JDBC property for the AllowSelfSignedCerts=1 property.

How to add the property allowing self-signed certificates:

  • At the bottom of the Connection details page under Connection settings select More options then JDBC properties.
  • Under JDBC properties configuration, select Add property then New property then enter AllowSelfSignedCerts as the key and 1 as the value.

When the AllowSelfSignedCerts property is set to 1, SSL verification is disabled. In this case, the connector does not verify the server certificate against the trust store, and does not verify if the server's host name matches the common name or subject alternative names in the server certificate.

The AllowSelfSignedCerts property and other JDBC properties are outlined in the Databricks driver documentation ↗. The JDBC properties outlined in this documentation are specific to the Databricks driver and will differ from other source types.

Virtual tables

Virtual tables allow you to connect to data registered in Databricks Unity Catalog. This allows you to both read and write to tables in Databricks from Foundry as well as push down compute to Databricks from pipelines in Foundry. This section provides additional details around using virtual tables with Databricks. This section is not applicable when syncing to Foundry datasets.

The Databricks connector offers enhanced functionality when using virtual tables to expose the features of Delta Lake and Apache Iceberg. This functionality requires external access to be enabled in Unity Catalog. When enabled, external access allows Foundry to access tables using the Unity REST API and Iceberg REST catalog, and read and write data in the underlying storage locations. Unity Catalog credential vending is used to ensure secure access to cloud object storage. In addition to enhanced functionality, this can also improve the performance of reads and writes against these tables.

The Databricks connector automatically exposes Delta Lake and Apache Iceberg functionality if you:

  1. Enable external access in Unity Catalog.
  2. Configure network egress policies that allow connectivity from Foundry to the table's storage location.
  3. Configure credentials on the source that can obtain vended credentials from Unity Catalog.

Connections to read or write tables will be made to the storage location directly using Delta Lake or Apache Iceberg clients. Databricks compute will not be used to read or write to the tables. The Unity Catalog REST APIs will be used for certain metadata operations such as determining the type of table being accessed.

Refer to the official Databricks documentation ↗ for more information on external access and how to determine the storage locations of tables. Refer to the Networking section of this documentation for details on enabling network access to storage locations.

Unity Catalog credential vending is required to use external access. Credential vending is not supported for all table types and table features. For example, views or tables with row filters do not support credential vending. Refer to the official Databricks documentation ↗ for more information on credential vending and the requirements.

If any of the above requirements are not met, connections to Databricks will be made using JDBC. JDBC is the same mechanism used for syncs. Refer to the official Databricks documentation ↗ for more information on JDBC connectivity to Databricks.

The table below highlights the virtual table capabilities that are supported for Databricks.

CapabilityStatus
Bulk registration🟢 Generally available
Automatic registration🟢 Generally available
Table inputs🟢 Generally available: tables, views, materialized views in Code Repositories, Pipeline Builder
Table outputs🟢 Generally available: Code Repositories, Pipeline Builder
Incremental pipelines🟢 Generally available [2]
Compute pushdown🟢 Generally available: Python transforms, Pipeline Builder

Consult the virtual tables documentation for details on the supported Foundry workflows where Databricks tables can be used as inputs or outputs. Functionality may vary depending on whether external access is enabled.

[2] To enable incremental support for Spark pipelines backed by Databricks virtual tables, external access must be enabled; incremental computation requires the ability to directly interact with Delta or Iceberg tables. Incremental compute on top of Delta tables relies on Change Data Feed ↗. Incremental compute on top of Iceberg tables relies on Incremental Reads ↗.

Table format and storage locations

The following table provides a summary of the supported formats and workflows when external access is or is not enabled.

Unity Catalog objectExternal access requiredFormatTable inputsTable outputs
Managed tableYesAvro ↗, Delta ↗, Parquet ↗✔️
Managed tableYesIceberg ↗✔️✔️
External tableYesDelta✔️✔️
External tableYesAvro, Parquet✔️
Managed tableNoTable ↗, View ↗, Materialized view✔️
External tableNoTable, view, materialized view✔️

When using Spark pipelines, virtual table outputs require external access to be enabled. The virtual table output must be either a managed Iceberg or external Delta table in Databricks.

Privileges on source credentials

For full feature support, we recommend providing the following privileges to the credentials provided for the source connection. These should be applied on either the catalog, schema, or table depending on the desired inheritance model.

CategoryPrivilegeNotes
PrerequisiteUSE CATALOG, USE SCHEMAMust be granted on the Databricks catalogs and schemas that will be used in Foundry.
MetadataBROWSERequired to explore source and register tables.
ReadSELECTRequired to read Databricks tables when using syncs or virtual table inputs.
EditMODIFYRequired to modify Databricks tables when using virtual table outputs.
CreateCREATE SCHEMA, CREATE TABLERequired to create Databricks tables when using virtual table outputs.
OtherEXTERNAL USE SCHEMAEnables external access to storage locations. Refer to external access for more details.

When using external tables, we recommend granting BROWSE, CREATE EXTERNAL TABLE, and EXTERNAL USE LOCATION privileges on the external locations being used. These are required when using virtual table outputs to create external tables.

Additionally, the credentials provided must have usage privileges on the warehouse or compute cluster provided in the source configuration.

Refer to the official Databricks documentation ↗ for more information on managing privileges in Unity Catalog.

Source configuration requirements

When using virtual tables, remember the following source configuration requirements:

  • You must use a Foundry worker source. Virtual tables do not support use of agent worker connections.
  • Ensure that bi-directional connectivity and allowlisting are established as described in the Networking section of this documentation, including the recommended networking to storage locations.
  • If using virtual tables in Code Repositories, refer to the Virtual Tables documentation for details of additional source configuration required.
  • You must specify a warehouse or compute cluster in the connection details using the HTTP path field. Refer to the official Databricks documentation ↗ on getting connection details for a Databricks compute resource.

See the Connection Details section above for more details.

Compute pushdown

Foundry offers the ability to push down compute to Databricks when using virtual tables. When using Databricks virtual tables registered to the same source as inputs and outputs to a pipeline, it is possible to fully federate compute to Databricks. See the Python documentation for details on how to push down compute to Databricks in Python Transforms. To push down compute to Databricks in Pipeline Builder, review the External pipelines documentation.

Use Databricks sources in code

You can use pro-code alternatives to connect to Databricks sources for more complex scenarios.

The example below demonstrates how to query Databricks from an external transform using a source configured with Workload identity federation, without storing a Databricks secret in Foundry. This requires a Foundry worker source and a federation policy in Databricks.

For a Databricks source, get_session_credentials().get().access_token returns the Foundry-issued OIDC identity token, not a Databricks access token. Databricks returns a 401 response if you send this token directly as a bearer token. The identity token must first be exchanged for a Databricks OAuth token using the OAuth 2.0 token exchange ↗ grant.

The Databricks SDK for Python ↗ (databricks-sdk) performs this exchange for you. Supply the Foundry identity token to the SDK through an IdTokenSource. The SDK exchanges it for a Databricks OAuth token and repeats the exchange whenever that token expires. Delegating the exchange matters because Foundry identity tokens expire after one hour and the Databricks token inherits that expiration. A long-running transform cannot exchange a token once at the start and reuse it throughout.

To pass a token directly to a client, read the exchanged token from the SDK configuration. The example below shows this approach with the Databricks SQL Connector for Python ↗ (databricks-sql-connector).

Read the workspace hostname, HTTP path, and service principal from source_configuration instead of setting them directly in your code. This property exposes the connection configuration of the source. This keeps the transform working if the source is later configured to use a different workspace.

The connection configuration API identifies Workload identity federation with the authentication type workflowIdentityFederation. Use that exact value when checking source_configuration.authentication.type, as shown below.

Both the token exchange and the query are HTTPS calls to your Databricks workspace, so the workspace hostname must be covered by the source's egress policies. Databricks sources do not expose an HTTPS connection, so the built-in HTTPS client returned by get_https_connection() is not available. Set the proxy and certificate environment variables from get_https_proxy_uri() and server_certificates_bundle_path before creating any client instead. These variables are required for sources that route egress through an agent proxy, and have no effect for sources that do not.

Example: Validate access and query a catalog

This example authenticates as the service principal, confirms the identity resolves, then runs a query using the exchanged token. Replace the query with your own to read application data.

The example passes a single access token to the SQL connector. This token is not refreshed during the connection's lifetime, so use this pattern only for short-lived queries that finish before the token expires. For long-running workloads that use Databricks REST APIs, use WorkspaceClient directly so the SDK can refresh credentials.

Copied!
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 import os import databricks.sql import polars as pl from databricks.sdk import WorkspaceClient, oidc from databricks.sdk.core import ( Config, credentials_strategy, oidc_credentials_provider, ) from transforms.api import Output, transform from transforms.external.systems import ( ResolvedSource, Source, external_systems, ) class FoundrySourceIdTokenSource(oidc.IdTokenSource): """Supply refreshable Foundry identity tokens to the Databricks SDK.""" def __init__(self, refreshable): self._refreshable = refreshable def id_token(self) -> oidc.IdToken: # This field holds the Foundry identity token, not a Databricks token. return oidc.IdToken(jwt=self._refreshable.get().access_token) def read_connection_settings(source: ResolvedSource): configuration = source.source_configuration if configuration.type != "databricks": raise ValueError("Expected a Databricks connection") authentication = configuration.authentication if authentication.type != "workflowIdentityFederation": raise ValueError("Expected OIDC federation authentication") client_id = authentication.service_principal_application_id if not client_id: raise ValueError("A service principal application ID is required") hostname = configuration.host_name.removeprefix("https://").rstrip("/") return hostname, configuration.http_path, client_id def configure_egress(source: ResolvedSource) -> None: proxy = source.get_https_proxy_uri() if proxy: os.environ["HTTPS_PROXY"] = proxy os.environ["HTTP_PROXY"] = proxy certificates = source.server_certificates_bundle_path if certificates: os.environ["REQUESTS_CA_BUNDLE"] = str(certificates) os.environ["SSL_CERT_FILE"] = str(certificates) def create_workspace_client( source: ResolvedSource, hostname: str, client_id: str, ) -> WorkspaceClient: refreshable = source.get_session_credentials() @credentials_strategy("foundry-source-oidc", []) def foundry_oidc(config: Config): return oidc_credentials_provider( config, FoundrySourceIdTokenSource(refreshable), ) return WorkspaceClient( host=f"https://{hostname}", client_id=client_id, credentials_strategy=foundry_oidc, ) def get_exchanged_access_token(client: WorkspaceClient) -> str: authorization = client.config.authenticate().get("Authorization", "") scheme, _, access_token = authorization.partition(" ") if scheme.lower() != "bearer" or not access_token: raise RuntimeError("Expected a Databricks bearer token") return access_token @external_systems( databricks_source=Source("<source_rid>"), ) @transform.using( output=Output("<output_dataset_rid>"), ) def compute(output, databricks_source: ResolvedSource): hostname, http_path, client_id = read_connection_settings( databricks_source ) configure_egress(databricks_source) client = create_workspace_client( databricks_source, hostname, client_id, ) # Validate authentication without logging or persisting identity details. client.current_user.me() # The SDK exchanges the Foundry identity token for a Databricks OAuth token. access_token = get_exchanged_access_token(client) with databricks.sql.connect( server_hostname=hostname, http_path=http_path, access_token=access_token, ) as connection: with connection.cursor() as cursor: cursor.execute(""" SELECT COUNT(*) AS visible_table_count FROM `<catalog_name>`.`information_schema`.`tables` WHERE table_schema = '<schema_name>' """) row = cursor.fetchone() if row is None: raise RuntimeError("Catalog query returned no result") visible_table_count = int(row[0]) output.write_table( pl.DataFrame({ "check_id": ["databricks_oidc_catalog_access"], "identity_check_succeeded": [True], "catalog_query_succeeded": [True], "visible_table_count": [visible_table_count], }) )

If you only need the Databricks REST APIs, use the WorkspaceClient directly and omit get_exchanged_access_token; the SDK attaches and refreshes the exchanged token for you on every call.

For more details on using session credentials with OIDC-enabled sources, review the Sources in Python documentation.

External models

Databricks models registered in Unity Catalog can be integrated to Foundry via:

Refer to the official Databricks documentation ↗ for more information on making models available in Unity Catalog, and to the guide on setting up Databricks external models in Foundry.