Skip to main content
Crusoe Support Help Center home page
Crusoe

How-To Deploy a Fine-Tuned Checkpoint or LoRA Adapter to a Self-Serve Deployment

Michael Xue
Michael Xue
Updated

Introduction

When a Serverless Fine-Tuning job succeeds, what you get is not a new model — it is a LoRA adapter. The base model's weights stay frozen during training, and the trainer learns a small set of low-rank matrices (tens to hundreds of MB) that sit on top of the base model at inference time. Every checkpoint the job emits is an independently usable adapter, registered in the same Intelligence Foundry model registry as the base models it was trained against.

Self-Serve Deployments are how you serve that adapter. A self-serve deployment reserves dedicated GPU capacity on Crusoe's managed inference engine and exposes an OpenAI-compatible Chat Completions endpoint. You pick the base model architecture, attach your checkpoint, choose an optimization profile, and set a replica count. Crusoe handles engine selection, tuning, and provisioning. Billing is per GPU-hour rather than per token, and throughput is bounded only by how many replicas you run — there is no shared rate limit.

The reason this matters: Serverless Inference, the multi-tenant option, serves base models only. It does not accept custom or fine-tuned models at all. If you fine-tuned a model on Crusoe and want to run inference against it on Crusoe, a self-serve deployment is the path — the checkpoint never leaves the platform, and there is no artifact to download, package, or upload to a serving stack.

This article walks through deploying a completed checkpoint from the Crusoe Cloud Console (the primary path), then covers the same operation from the command line using the crusoe CLI for project lookup and curl against the self-serve deployments API.

Prerequisites

  • Serverless Fine-Tuning Job in succeeded State With at Least One Checkpoint
  • Access to the Intelligence Foundry App in the Crusoe Cloud Console
  • Intelligence API Key
  • Crusoe CLI Installed and Configured (Command Line Path Only)
  • Self-Serve Deployment Quota in the Target Project

ℹ️ Note: A checkpoint can only be deployed onto the base model architecture it was trained against. A LoRA adapter trained on meta-llama/Llama-3.1-8B-Instruct will not appear as a selectable checkpoint under any other base model. See Available models for the base models Self-Serve Deployments supports with LoRA adapters.

Instructions

Step 1: Pick the Checkpoint You Want to Serve

A fine-tuning job saves a checkpoint every checkpoint_steps training steps (default 100) plus a final one when training ends, so a finished job usually has several. A short run that finishes at step 109, for example, yields checkpoint-100-<job_id> and checkpoint-109-<job_id>. They are not interchangeable — later is not automatically better, particularly if the run started to overfit.

  1. Sign in to the Crusoe Cloud Console and switch to the Intelligence Foundry app using the app switcher in the bottom-left corner.
  2. In the left navigation, select Model Shaping > Serverless Fine-Tuning.
  3. Click the row for your completed job to open its details page. The page shows the training and validation loss curves and the list of saved checkpoints.
  4. Compare loss across checkpoints and note which one you want to deploy. If you supplied a validation dataset, lowest validation loss is the usual pick. If you did not, you may only have training loss to go on — and training loss alone cannot tell you whether a later checkpoint has started to overfit.Fine-tuning job details page showing loss curves and saved checkpoints

💡 Tip: Supply a validation dataset when you launch the job, even a small one. Without it, checkpoints can come back with valid_loss: null, and you are choosing between checkpoints on training loss alone.

Step 2: Deploy From the Console

There are two console entry points that land on the same Create deployment form. Pick whichever matches where you are.

Option A — From the fine-tuning job details page (fastest)

  1. On the job details page from Step 1, locate the checkpoint you selected.
  2. Click the three-dot menu on that checkpoint's row and select Deploy.
  3. The Create deployment form opens with the base model and checkpoint pre-populated. Continue to Step 3.Checkpoint row menu with the Deploy option

Option B — From the Self-Serve Deployments page

  1. In the left navigation, select Inference > Self-Serve Deployments. The table lists every deployment in the project with its status, model, hardware, replica count, and metadata.
  2. Click Create deployment.
  3. Select the base model that matches your checkpoint's architecture.
  4. Select your fine-tuned checkpoint from the list that appears under the base model. Only checkpoints trained on that base model are listed. Continue to Step 3.Create deployment form with base model and fine-tuned checkpoint selected

Step 3: Choose an Optimization Profile and Replica Count

The rest of the form is the same regardless of how you arrived:

  1. Select a deployment configuration. This is a per-deployment choice that determines how the inference engine is tuned:
    • Responsiveness — optimized for low latency and time-to-first-token. Use for interactive, real-time, or latency-sensitive applications.
    • Throughput — optimized for token volume and cost-per-token. Use for batch processing and high-volume pipelines.
    • Balanced — hybrid tuning for general-purpose production traffic with moderate volume and latency requirements.
  2. Set the replica count. Throughput scales with replicas, and so does cost. The hourly rate for your chosen configuration is displayed on the form before you confirm.
  3. Give the deployment a name. This becomes the deployment alias, which is the value you pass as model in inference requests.
  4. Click Create.Create deployment form showing configuration, replica count, and deployment name

⚠️ Warning: Billing starts when the deployment is created, not when it first serves traffic. You are paying GPU-hours for reserved capacity through the Creating state, and billing stops only when the deployment is deleted.

Step 4: Wait for the Deployment to Reach Ready

Provisioning reserved capacity can take up to 40 minutes. The deployment appears in the Self-Serve Deployments table immediately in the Creating state. Click the row to watch the activity log and progress.

State Meaning
Creating Capacity is being provisioned and the engine is starting up
Ready Serving traffic
Scaling up / Scaling down Replica count is changing; traffic is still served
Syncing Alias is being updated
Failed Creation or update did not complete
Deleting / Deleted Teardown in progress or complete

If the deployment lands in Failed, open the row and check the activity log. The most common causes are insufficient self-serve quota for the requested replica count, a checkpoint/base-model mismatch, or (via the API) a flavor_id that is not lora_suitable — see Step 6d. If the log does not make the cause clear, open a support ticket with the deployment ID.Self-Serve Deployments table with a deployment in the Creating state

By default, console notifications are sent when a self-serve deployment is created, deleted, or scaled. To route them to Slack or a webhook, click the bell icon in the top-right corner of the console, select All Notifications, and click Manage Slack/Webhook.

Step 5: Send a Test Request

Once the status is Ready, the deployment serves the standard OpenAI-compatible Chat Completions API. Authenticate with an Intelligence API key and pass the deployment alias as model:

export API_TOKEN='<YOUR_INTELLIGENCE_API_KEY>'

curl 'https://api.inference.crusoecloud.com/v1/chat/completions' \
  --request POST \
  --header 'Content-Type: application/json' \
  --header "Authorization: Bearer $API_TOKEN" \
  --data '{
    "model": "<YOUR_DEPLOYMENT_ALIAS>",
    "messages": [
      {"role": "system", "content": "You are a helpful assistant."},
      {"role": "user", "content": "Which countries are represented?"}
    ]
  }'

The response should reflect your fine-tuned behavior, not the base model's. If you trained a classifier, you should get a label back; if you trained for a persona, the tone should be recognizably different from the base model. A quick sanity check against the same prompt on the base model via Serverless Inference is the fastest way to confirm the adapter is actually being applied.

The deployment's detail page in the console also includes ready-to-copy sample code (Python, TypeScript, cURL) pre-filled with your alias.

Step 6 (Alternative): Deploy From the Command Line

The published crusoe CLI does not currently include a subcommand for self-serve deployments or fine-tuning jobs — its scope is compute, networking, storage, and project management. The command-line path therefore uses the crusoe CLI to resolve your project ID and curl against the self-serve deployments REST API to create and manage the deployment.

6a. Resolve your project ID with the Crusoe CLI

The self-serve deployments API is project-scoped. If you don't already have the project ID handy:

crusoe whoami
crusoe projects list

Copy the ID of the project that owns your fine-tuning job.

6b. Set up your shell

Both the Managed AI API (fine-tuning, files) and the Cloud API (self-serve deployments) accept the same Intelligence API key:

export API_TOKEN='<YOUR_INTELLIGENCE_API_KEY>'
export PROJECT_ID='<YOUR_PROJECT_ID>'
export FT_URL='https://api.intelligence.crusoecloud.com'
export DEPLOYMENT_URL="https://api.crusoecloud.com/v1/projects/${PROJECT_ID}/foundry/selfserve"

ℹ️ Note: Crusoe's API endpoints sit behind Cloudflare, which rejects the default User-Agent of some Python HTTP libraries (observed with urllib) with HTTP 403 / Error 1010: Access denied. curl works out of the box. If you script against these APIs, set an explicit User-Agent header (for example curl/8.7.1) or the request never reaches Crusoe — the error looks like an auth failure but has nothing to do with your key.

6c. Retrieve the fine-tuned model identifier

Every checkpoint has its own fine_tuned_model_id in the form ftmodel-<uuid>. That is the value the self-serve deployments API expects — not the checkpoint name, and not the adapter:... string in the job's result_files. List the job's checkpoints to get it:

export JOB_ID='ftjob-<YOUR_JOB_ID>'

curl "$FT_URL/v1/fine_tuning/jobs/${JOB_ID}/checkpoints" \
  --header 'Accept: application/json' \
  --header "Authorization: Bearer $API_TOKEN"

Each entry in data is one checkpoint:

{
  "data": [
    {
      "id": "5e1c2d3a-4b5c-4d6e-8f70-1a2b3c4d5e6f",
      "fine_tuned_model_checkpoint": "checkpoint-109-ftjob-0f3a9c2e7b1d4e8fa6c5d4b3a2910fed",
      "step_number": 109,
      "metrics": {
        "train_loss": 0.5037,
        "valid_loss": null
      },
      "fine_tuning_job_id": "ftjob-0f3a9c2e7b1d4e8fa6c5d4b3a2910fed",
      "fine_tuned_model_id": "ftmodel-1a2b3c4d-5e6f-4a7b-8c9d-0e1f2a3b4c5d"
    },
    {
      "fine_tuned_model_checkpoint": "checkpoint-100-ftjob-0f3a9c2e7b1d4e8fa6c5d4b3a2910fed",
      "step_number": 100,
      "metrics": { "train_loss": 0.5229, "valid_loss": null },
      "fine_tuned_model_id": "ftmodel-9f8e7d6c-5b4a-4c3d-9e2f-1a0b9c8d7e6f"
    }
  ]
}
  • fine_tuned_model_checkpoint — the checkpoint name as it appears in the console (checkpoint-<step>-<job_id>). Use this to match what you picked in Step 1.
  • fine_tuned_model_id — the ftmodel-... identifier to pass to the deployments API. Note that it differs per checkpoint.
  • metrics.train_loss / metrics.valid_loss — use these to pick the checkpoint. valid_loss is null if the job ran without an evaluation pass.

ℹ️ Note: The job object itself (GET /v1/fine_tuning/jobs/{id}) also carries a top-level fine_tuned_model field, but it points at only one checkpoint (the final one). Its result_files entries (adapter:checkpoint-<step>-<job_id>:<project_id>:<hash>) are Files API handles for downloading the adapter, not deployment identifiers. Use the checkpoints list when you want a specific checkpoint.

6d. List available flavors

A "flavor" is the API's term for a base model + optimization profile + hardware combination — what the console calls a deployment configuration. List the flavors available to your project:

curl "$DEPLOYMENT_URL/flavors" \
  --header "Authorization: Bearer $API_TOKEN"

Each flavor looks like this:

{
  "id": "c0ffee00-1234-4abc-9def-0123456789ab",
  "model_name": "Qwen/Qwen3.5-2B",
  "flavor_type": "throughput",
  "gpu_type": "h100",
  "gpus_per_instance": 1,
  "quantization": "bf16",
  "lora_suitable": true,
  "provider": "Qwen",
  "price_per_hour": "5.5",
  "context_length": 262144
}
  • id — the value to pass as flavor_id.
  • model_name — must match your checkpoint's base model architecture.
  • flavor_type — responsiveness, throughput, or balanced.
  • lora_suitable — whether this flavor can host a LoRA adapter. This is the filter that matters.
  • price_per_hour — per-replica hourly cost in USD.

⚠️ Warning: Not every flavor for a given base model accepts LoRA adapters. Some (model, profile, GPU) combinations — typically the FP8-quantized responsiveness tiers — have lora_suitable: false. Filter on lora_suitable: true before picking a flavor_id, and be aware that some base models have only one LoRA-capable flavor, which means the optimization profile is effectively chosen for you.

To list only the flavors that can host your checkpoint:

curl -s "$DEPLOYMENT_URL/flavors" \
  --header "Authorization: Bearer $API_TOKEN" \
  | jq '.flavors[] | select(.lora_suitable and .model_name == "<YOUR_BASE_MODEL>")'

6e. Create the deployment

curl "$DEPLOYMENT_URL/deployments" \
  -X POST \
  -H "Authorization: Bearer $API_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "fine_tuned_model_id": "<YOUR_FINE_TUNED_MODEL_ID>",
    "flavor_id": "<YOUR_FLAVOR_ID>",
    "deployment_name": "<YOUR_DEPLOYMENT_ALIAS>",
    "replicas": 1
  }'
  • fine_tuned_model_id — the ftmodel-... identifier from step 6c. Omit this field to deploy the base model alone.
  • flavor_id — the flavor id from step 6d. It must be lora_suitable and its model_name must match the checkpoint's base model.
  • deployment_name — the alias you will pass as model in inference requests. Must be unique within the project.
  • replicas — integer replica count. Must fit within your project's self-serve quota.

6f. Poll for ready

curl "$DEPLOYMENT_URL/deployments" \
  -X GET \
  -H "Authorization: Bearer $API_TOKEN"

The response is a deployments array. A deployment serving a fine-tuned checkpoint looks like this:

{
  "deployments": [
    {
      "id": "qwen-qwen3-5-2b-a1b2c3d4",
      "deployment_name": "my-finetune-v1",
      "status": "ready",
      "model": "Qwen/Qwen3.5-2B",
      "base_model": "Qwen/Qwen3.5-2B",
      "quantization": "bf16",
      "gpu_type": "h100",
      "gpus_per_replica": 1,
      "target_replicas": 1,
      "available_replicas": 1,
      "fine_tuned_model": "ftmodel-1a2b3c4d-5e6f-4a7b-8c9d-0e1f2a3b4c5d",
      "created_at": "2026-09-15T14:02:11Z",
      "context_length": 262144
    }
  ]
}
  • id — the deployment ID used in the per-deployment endpoints below. It is derived from the base model name plus a short hash, not a UUID.
  • status — lowercase in the API (creating, ready, ...) even though the console capitalizes it.
  • fine_tuned_model — confirms which checkpoint the deployment is serving. If this is empty, you deployed the base model without the adapter.
  • target_replicas / available_replicas — when these differ, the deployment is still scaling.

Or fetch a single deployment by ID for its activity log and metadata:

curl "$DEPLOYMENT_URL/deployments/<YOUR_DEPLOYMENT_ID>" \
  -X GET \
  -H "Authorization: Bearer $API_TOKEN"

When status is ready, send the test request from Step 5.

6g. Scale, rename, or tear down

Update the replica count (the deployment keeps serving during the change):

curl "$DEPLOYMENT_URL/deployments/<YOUR_DEPLOYMENT_ID>" \
  -X PATCH \
  -H "Authorization: Bearer $API_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"replicas": 2}'

Rename the alias (status goes to Syncing, then back to Ready):

curl "$DEPLOYMENT_URL/deployments/<YOUR_DEPLOYMENT_ID>" \
  -X PATCH \
  -H "Authorization: Bearer $API_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"deployment_name": "<NEW_ALIAS>"}'

Delete the deployment and stop billing:

curl "$DEPLOYMENT_URL/deployments/<YOUR_DEPLOYMENT_ID>" \
  -X DELETE \
  -H "Authorization: Bearer $API_TOKEN"

Example

A platform team fine-tunes meta-llama/Llama-3.1-8B-Instruct as a support-ticket classifier: the system prompt lists the allowed categories and every training example ends in an assistant turn containing a single label such as billing or gpu_fault. They supply a small held-out validation file, and the job finishes with checkpoints at steps 100, 200, and 250.

On the job details page, validation loss bottoms out at step 200 and rises slightly at step 250, so they deploy checkpoint-200-<job_id> rather than the final checkpoint. They open the checkpoint's three-dot menu, select Deploy, choose the Throughput configuration because tickets are classified in batches by a queue worker, set one replica, and name the deployment ticket-classifier-v1.

Once the deployment reaches Ready, the queue worker sends each ticket body to the Chat Completions endpoint with "model": "ticket-classifier-v1" and receives a bare category label back. When a later training run produces a better adapter, they deploy it as ticket-classifier-v2, move the worker over, and delete the v1 deployment to stop its GPU-hour billing.

Related Articles

Additional Resources

Related to

Attachments

Was this article helpful?

0 out of 0 found this helpful

Still need help?

Our support team is ready to assist you with any questions.

Have more questions? Submit a request

Related Articles

Recently Viewed

Comments

0 comments

Article is closed for comments.