Skip to main content
Crusoe Support Help Center home page
Crusoe

FAQ: Common dataset errors in Serverless Fine-Tuning

Michael Xue
Michael Xue
Updated

Introduction

Most Serverless Fine-Tuning jobs that fail do so in the first few seconds, in validating_files, before a GPU is ever scheduled. The cause is almost always the training file, and the fix is almost always on your side.

The platform accepts a single training file in JSONL or Parquet format, 3 GB or smaller, in which each example is a chat conversation in a messages array. This FAQ covers JSONL, the format most export pipelines produce.

This FAQ answers the dataset questions that come up most often, in roughly the order you will hit them: upload, job creation, validation, and the quiet problems that pass validation but hurt the adapter.

Each answer below shows the offending JSONL line as it appears in a real file, next to the corrected line, so you can pattern-match against your own data. Where an answer mentions "the local validator", it means the Python script in the companion article Serverless Fine-Tuning job fails with dataset format or validation errors (Step 4), which reproduces the platform's checks on your machine before you upload.

ℹ️ Note: For the full troubleshooting procedure, including Console and curl steps, see Serverless Fine-Tuning job fails with dataset format or validation errors. This FAQ is the short version.

Question 1: My file is valid JSON. Why is it rejected?

Answer
Because JSONL is not JSON. A .json file holding one array of examples is a single JSON document; the platform wants one JSON object per line. Renaming the extension does not convert it. Loop over the array and write json.dumps(ex) + "\n" for each element. The local validator's tell is line 1: top-level value must be an object, got list.

Rejected — the whole file is one JSON array:

[{"messages": [...]}, {"messages": [...]}, {"messages": [...]}]

Accepted — one object per line, no commas between lines, no surrounding brackets:

{"messages": [...]}
{"messages": [...]}
{"messages": [...]}

Question 2: The upload succeeded but the job failed seconds later with an error on training_file. Which one is wrong?

Answer
Both are behaving correctly. Upload checks only the extension (.jsonl or .parquet), the 3 GB limit, and purpose=fine-tune. Content is checked when the job enters validating_files. A fast failure with trained_tokens: null and finished_at close to created_at is a dataset failure, not a training failure. Read the error object on the job; for a schema failure the message quotes the offending line, for example Line doesn't contain 'messages' key: {'input': '...', 'output': ''}.

Question 3: How many kinds of mistakes are there, and in what order should I fix them?

Answer
The ones we see most often are nine: a UTF-8 BOM, CRLF line endings, a trailing comma, a prompt/completion row, a row with no assistant turn, a blank line, the role human, an empty assistant reply, and a literal newline inside a string. They are not independent (a stray \r can hide a schema problem, which hides a semantic one), so repair in layers, from the outside in:

  1. Normalize the bytes: strip the BOM, convert CRLF to LF, drop blank lines.
  2. Make every line parse: strict json.loads, then retry with trailing commas removed, then join with following lines until it parses; drop and report anything left.
  3. Normalize the schema: map prompt/completion, instruction/output, and ShareGPT conversations onto messages; rename human to user and gpt/bot/model to assistant.
  4. Apply the semantic checks: non-empty messages, standard roles, content on every non-assistant turn, at least one assistant turn with content or tool_calls. The platform enforces only the first of these at validation time; the rest surface later as tokenization errors or as a poor adapter, so check them yourself.
  5. Separate repairs from drops: layers 1 to 3 lose no data; layer 4 failures are content the file does not contain, so drop the row and list it for a human.
  6. Re-validate strictly and exit non-zero on any residue, so the check can run in your export pipeline ahead of the upload step.

Layers 1 to 3 can be scripted safely; layer 4 is where you decide what to drop.

Question 4: I use prompt/completion like the OpenAI legacy format. Is that supported?

Answer
No. Only the chat format with a messages array is accepted. Map prompt to a user turn and completion to the assistant turn. The same applies to Alpaca instruction/output and ShareGPT conversations with from/value.

Rejected (Line doesn't contain 'messages' key):

{"prompt": "Why is my bill higher than the estimate?", "completion": "billing"}

Accepted:

{"messages": [{"role": "user", "content": "Why is my bill higher than the estimate?"}, {"role": "assistant", "content": "billing"}]}

Question 5: Do I need a system turn in every example?

Answer
No. system is optional. What is required is at least one assistant turn with content (or tool_calls), because loss is computed only on assistant tokens. A row with only system and user turns has nothing to learn from; the local validator reports it as no assistant turn, and on the platform it contributes no training signal at best and fails tokenization at worst.

Rejected — the conversation ends before anyone answered:

{"messages": [{"role": "system", "content": "Classify the ticket: billing, gpu_fault, networking, quota, or other."}, {"role": "user", "content": "Getting 'quota exceeded' when I launch a second 8xH200 VM."}]}

Accepted:

{"messages": [{"role": "system", "content": "Classify the ticket: billing, gpu_fault, networking, quota, or other."}, {"role": "user", "content": "Getting 'quota exceeded' when I launch a second 8xH200 VM."}, {"role": "assistant", "content": "quota"}]}

Question 6: Which roles are valid?

Answer
Use system, user, assistant, and tool. The platform's validation step checks that each line is valid JSON with a messages key; it does not enforce a role list. Roles are interpreted later by the base model's chat template, so a non-standard role such as human or gpt is not reliably rejected: depending on the model it either errors during tokenization with a less helpful message, or is rendered as a literal role string and trains the model on a format it will never see at inference. Rename before uploading (human to user; gpt, bot, model to assistant). The local validator reports these as unknown role so you catch them before upload. tool turns and tool_calls on assistant turns are supported through the chat template; make sure every tool_calls[].function.arguments value is a valid JSON string, and keep to base models whose template supports tool calling.

Wrong — ShareGPT-style role names:

{"messages": [{"role": "human", "content": "Invoice shows $0 for storage but I have 20 TB provisioned."}, {"role": "gpt", "content": "billing"}]}

Correct:

{"messages": [{"role": "user", "content": "Invoice shows $0 for storage but I have 20 TB provisioned."}, {"role": "assistant", "content": "billing"}]}

Question 7: The local validator says assistant turn has neither content nor tool_calls, but I can see the text in my source data. What happened?

Answer
The export wrote "" or null for that row, usually because the label or reply column was empty in the source system. An assistant turn may have empty content only when it carries tool_calls. Restore the value from the source or drop the example; do not fill it with a placeholder.

Rejected — nothing to learn from:

{"messages": [{"role": "user", "content": "Training job dies with ECC uncorrectable error on GPU 3."}, {"role": "assistant", "content": ""}]}

The one legitimate empty content is on a turn that carries tool_calls:

{"role": "assistant", "content": "", "tool_calls": [{"id": "call_1", "type": "function", "function": {"name": "lookup_gpu", "arguments": "{\"gpu\": 3}"}}]}

Question 8: My ticket bodies contain newlines and quotes. How do I include them?

Answer
Let json.dumps escape them: a newline becomes the two characters \n, a quote becomes \". If you build lines by string formatting or concatenation, a literal newline splits one example across several physical lines and every one of them fails to parse. This is the single most common failure in datasets built from real support text.

Rejected — one example spread over three physical lines (Unterminated string on the first, Expecting value on the next two):

{"messages": [{"role": "user", "content": "Job crashed with:
CUDA error: uncorrectable ECC
please help"}, {"role": "assistant", "content": "gpu_fault"}]}

Accepted — the newlines are the two characters \n inside the string, which json.dumps writes for you:

{"messages": [{"role": "user", "content": "Job crashed with:\nCUDA error: uncorrectable ECC\nplease help"}, {"role": "assistant", "content": "gpu_fault"}]}

Question 9: Can I have several assistant turns in one example?

Answer
Yes for ordinary multi-turn chat. Not yet if those assistant turns carry reasoning content such as a thinking field; multi-turn reasoning is unsupported. Strip the reasoning fields or split the conversation into single-turn examples. If you need multi-turn reasoning, open a support ticket.

Not supported yet — two assistant turns, both with thinking:

{"messages": [{"role": "user", "content": "What is 2+2?"}, {"role": "assistant", "content": "4", "thinking": "Add them."}, {"role": "user", "content": "And 3+3?"}, {"role": "assistant", "content": "6", "thinking": "Add them."}]}

Either works — single-turn with reasoning, or multi-turn without it:

{"messages": [{"role": "user", "content": "What is 2+2?"}, {"role": "assistant", "content": "4", "thinking": "Add them."}]}
{"messages": [{"role": "user", "content": "What is 2+2?"}, {"role": "assistant", "content": "4"}, {"role": "user", "content": "And 3+3?"}, {"role": "assistant", "content": "6"}]}

Question 10: The file has a BOM or Windows line endings. Does that matter?

Answer
It can. Some parsers reject a BOM outright (Unexpected UTF-8 BOM), and a trailing \r can land inside the last string value or fail parsing. Detect with file train.jsonl or head -c 3 train.jsonl | xxd; fix by rewriting with encoding='utf-8' and newline='\n'. The local validator tolerates both deliberately so it can report the errors underneath.

What the two problems look like on disk; neither is visible in an editor:

$ file train.jsonl
train.jsonl: Unicode text, UTF-8 (with BOM) text, with CRLF line terminators

$ head -c 3 train.jsonl | xxd
00000000: efbb bf                                  ...

A parser that does not strip the BOM reports the first line as Unexpected UTF-8 BOM (decode using utf-8-sig): line 1 column 1.

Question 11: Why does the validator pass locally but the job still fails in validating_files?

Answer
Three possibilities, in order of likelihood. First, you validated a different file from the one you uploaded; re-download it with GET /v1/files/<id>/content and validate that. Second, a row exceeds the model's maximum sequence length, which a local validator cannot check without the tokenizer; resubmit with "overlong_row_behavior": "drop" or split the long conversation. Third, individual rows are very large: keep each example (one JSONL line) under roughly 100 KB or 32k tokens, and split long agentic or multi-turn sessions into several shorter examples. Rows longer than the model's context window cannot be trained on in any case. If none applies, open a support ticket with the job ID, file ID, model ID, the error object, and the validator output.

Question 12: Can I fix the file and retry the same job?

Answer
No. The Files API has no update operation and the Fine-tuning API has no retry: the only operations on an existing job are retrieve, list checkpoints, list events, get metrics, and cancel. Upload the corrected file as a new file, take the new ID from the response, and create a new job pointing at it. Delete the rejected file afterwards to keep the file list clean.

Question 13: The job passed validation but the adapter is bad. Is that a dataset problem?

Answer
Often. Two patterns pass validation but degrade training: gpt-oss assistant turns without "thinking": "" (the model fights its pretrained analysis-then-final format and emits stray reasoning text), and llama-3 assistant turns with both content and tool_calls (the content is never rendered). The local validator reports these as WARN, not ERROR, when MODEL is set to your base model. Also confirm you did not pass the same file as both training_file and validation_file; validation loss will track training loss and tell you nothing.

Question 14: How big should the dataset be?

Answer
Smaller and cleaner beats larger and noisy for LoRA. A few thousand consistent examples is a normal starting point. The 3 GB limit is a ceiling, not a target, and a job takes exactly one training_file, so splitting a file across uploads does not work around it.

Related Articles

Additional Resources

Related to

Was this article helpful?

0 out of 0 found this helpful

Still need help?

Our support team is ready to assist you with any questions.

Have more questions? Submit a request

Related Articles

Recently Viewed

Comments

0 comments

Article is closed for comments.