Docs
LoginTry for free

Get started

Build

Platform

Data & content

Automation

Developers

Platform

Custom models

Use DeepSeek or another OpenAI-compatible provider in Turbofy’s built-in chat. You can connect a hosted service or a model running on your own computer.

How it works

Add your provider’s API address, API key and model ID to your workspace. Its models will then appear in the chat model picker.

Providers are managed in Settings → Chat models. Each provider can have several models, with a display name for each. The steps below use DeepSeek as an example.

Custom-provider tokens do not consume Turbofy credits. Your provider bills API usage separately. Additional data Turbofy creates to keep the chat functional still consumes credits.

Add a third-party provider

  1. 1Get an API key, the OpenAI-compatible base URL and the exact model ID from your provider.
  2. 2Create a workspace secret containing the provider’s API key. Give it a recognizable name, such as DeepSeek API key.
  3. 3Open workspace Settings. Under Chat models, click New provider and choose Custom.
  4. 4Enter the provider’s base URL in Endpoint, including /v1 if required by that provider. Do not enter the full /chat/completions URL.
  5. 5Select your secret in API key secret. Enter the model ID and the name you want to see in the chat picker. Use Add model if you want to add more models from this provider.
  6. 6Click Save provider.
example: DeepSeek
Provider:       Custom
Endpoint:       https://api.deepseek.com
API key secret: DeepSeek API key
Model ID:       deepseek-flash
Display name:   DeepSeek Flash

Use deepseek-flash as the model ID. DeepSeek Flash is the display name shown in the chat picker.

Choose it in the chat

Open the model picker at the bottom of the chat. Under Workspace models, choose the name you saved, such as DeepSeek Flash, and send a test message.

Auto uses Turbofy’s model selection. Choose your workspace model by name to use your custom provider.

Thinking and output

Expand Thinking and output in the provider editor to set a model’s thinking options and output limit. These settings can override the global output limit. Global instructions still apply.

Check the provider’s documentation for supported thinking options and output limits.

A model’s context window includes both the prompt and generated output. Its maximum output length may be smaller than its total context window.

Troubleshooting custom providers

  • Model missing from the picker: confirm that you saved the provider in the current workspace, then look under Workspace models.
  • Authentication error: check that you selected the right secret and that it contains a valid API key.
  • Model not found: use the exact model ID supplied by the provider or returned by its models endpoint. The friendly display name can differ.
  • Connection failure: check the provider’s base URL, service availability and any account limits.
  • Unsupported thinking or output option: choose settings accepted by the provider and model.
  • Tool-use errors: a model may handle ordinary chat but struggle with Turbofy tasks that use tools. Test both before relying on it.

Optional: use local AI models

To run a model on your computer, choose one of the apps below. Follow its instructions to download a model and start its OpenAI-compatible API server. Turbofy uses that server to send messages to the model.

Desktop apps

  • LM Studio

    A desktop app for downloading and running models. Its Developer tab lets you start a local API server.

  • Llama (llama.cpp)

    The llama.cpp team’s desktop app for Mac and Windows, with model downloads, browser chat and a local OpenAI-compatible API.

  • mlx-dspark

    A native Mac app for Apple Silicon, with a model manager, chat and an OpenAI-compatible server.

Engines for more technical setups

  • llama.cpp server documentation

    Set up the llama.cpp server directly if you prefer command-line controls.

  • vLLM documentation

    An inference engine with an OpenAI-compatible server.

  • ExLlamaV3

    An inference engine for running models on consumer GPUs.

  • TabbyAPI documentation

    The OpenAI-compatible API server for ExLlama.

Choose a model that fits your computer and test it in the app first. Note its model ID and local server address. Your computer must stay awake and the server must stay running while you use it from Turbofy.

Turbofy cannot connect to localhost on your computer. It needs an HTTPS address it can reach over the internet. Tailscale Funnel provides one, as shown below. Messages still pass through Turbofy, and the credit rules above still apply.

Connect your local model with Tailscale

Tailscale connects your devices through a private network. Its Funnel feature gives your local model server a public HTTPS address that Turbofy can use.

  • Tailscale Funnel setup guide

    Install Tailscale on the computer running your model, sign in, and follow the setup guide to enable Funnel.

Anyone on the internet can reach a Funnel address. Set your model server to require an API key or token before turning Funnel on.

Example with LM Studio

  1. 1In LM Studio, download and load a model, then open the Developer tab and start the server. This example assumes the local address is http://127.0.0.1:1234.
  2. 2In Server Settings, enable authentication and create an API token. Save that token as a Turbofy workspace secret, such as Local model API key.
  3. 3With Tailscale installed and signed in on the same computer, open Terminal or PowerShell and run the command below. Follow any setup link Tailscale displays.
  • LM Studio API-token instructions

    How to enable authentication and create a token in Server Settings.

terminal
tailscale funnel --bg --https=8443 http://127.0.0.1:1234

Replace 1234 with the port shown by your model app if it differs. 8443 is the public HTTPS port, and --bg keeps the tunnel running after you close the terminal. If Linux reports a permission error, run the same command with sudo at the beginning.

Copy the HTTPS address shown by Tailscale and add /v1 for the OpenAI-compatible endpoint. For example:

example endpoint
https://YOUR-COMPUTER.YOUR-TAILNET.ts.net:8443/v1

In Turbofy, return to Settings → Chat models → New provider. Choose Custom, paste this endpoint, select your API-key secret, and enter the exact model ID from your model app. Save it, then select its display name under Workspace models in the chat picker.

Keep your computer, model server and Tailscale running. If the connection stops working, check the tunnel with:

terminal
tailscale funnel status

To stop sharing this endpoint:

terminal
tailscale funnel --bg --https=8443 off
  • Funnel command reference

    Command options and troubleshooting. Use Funnel for a public address. Tailscale Serve only shares within your private network.

On this page