Introduction
Crusoe Managed Inference serves every model in its catalog behind one OpenAI-compatible endpoint, https://api.inference.crusoecloud.com/v1. The request shape is the same for all of them — /v1/chat/completions, a messages array, the OpenAI SDK you already have installed. What differs is what each model can take in and what it can do, and that is what this article is about.
There are three capability tiers in the catalog, and they build on each other.
Text is the baseline. Every chat model accepts text messages and returns text. This is system / user / assistant turns, streaming, temperature, max_tokens — the territory covered by the Getting Started article.
Image input is a per-model property, not an endpoint property. Models tagged image text to text in the catalog — Nemotron 3 Nano Omni, Gemma 4 31B, Kimi K2.6, GLM-5.3 Flash, and Qwen3.8 27B at the time of writing — accept an image_url content part in the same user message as your text and reason over the pixels. There is no separate vision endpoint. You send the image to /v1/chat/completions and the model either understands it or it doesn't, which is why Step 1 exists.
Browser control is not a fourth kind of API call. It is image input plus tool calling, run in a loop: your code screenshots a browser, sends the screenshot and the task to the model, gets back a tool_calls entry saying where to click or what to type, executes it with Playwright, screenshots again, and repeats until the model answers in plain text. The model never touches the browser. It plans; your harness has the hands.
Which model you put in that loop matters more than the loop itself. The catalog has one built for the job — Yutori's Navigator n2 (yutori/n2, tagged browser use / computer use) — and several general vision models that can be pressed into it.
n2 is a 27B model trained to operate a desktop. Hand it a screenshot and a computer_batch tool definition and it returns an ordered batch of GUI primitives in a fixed vocabulary, with every coordinate in a normalized 1000×1000 space no matter what your schema says. That consistency is the point: one coordinate convention, one executor, and it holds across tasks. The model page in the console lists a 65.2% partial score on OSWorld 2.0. Step 4 covers it.
General vision models with tools support — Nemotron 3 Nano Omni, Gemma 4, GLM-5.3 Flash — will also emit click(x, y) calls against a screenshot, and Step 5 shows that pattern for when you want an open-weights model or one that doubles as your general assistant. The catch is that each has its own idea of what a coordinate is, and none of them was trained to your schema. Your executor has to detect the convention before it can click.
One thing to know before Step 4: n2 has no built-in tools on this endpoint. It is served like every other model in the catalog — the same /v1/chat/completions contract, the same request fields — and nothing is injected on its behalf. Send yutori/n2 a screenshot with no tools and you get a confident paragraph describing a click that never happened. You pass the computer_batch definition yourself.
Cost-wise, every turn of the loop resends the same task and tool schema plus a growing history. Crusoe's inference engine with MemoryAlloy routes requests cache-aware, so the unchanged prefix of each request is served from cache at the cached rate on the catalog card — $0.05 versus $0.50 per million input tokens for n2. usage.prompt_tokens_details.cached_tokens on each response tells you how much of the prompt hit cache; commonly more than half of it after the first turn.
Prerequisites
- Crusoe Account With Intelligence Foundry Access
- Inference API Key Generated From the Intelligence Foundry in the Crusoe Cloud Console
- Python 3.10+ With the
openaiLibrary Installed in a Virtual Environment (python3 -m venv .venv && source .venv/bin/activate && pip install openai) -
curlandjqInstalled (Step 1 Only) - Playwright Installed (
pip install playwright && playwright install chromium) (Browser Control Only) - Pillow Installed (
pip install pillow) (Step 5 Coordinate Calibration Only)
ℹ️ Note: Crusoe API keys can contain
$characters. Export the key wrapped in single quotes (export CRUSOE_API_KEY='...') — double quotes let Bash/Zsh expand segments like$2aas variables, mangling the key and producing401 Unauthorized. Safer still, read it from stdin so it never touches the command line or your shell history:read -rs CRUSOE_API_KEY && export CRUSOE_API_KEY, then paste the key and press Enter.
Instructions
Step 1: Check Which Capabilities a Model Actually Has
Do not guess from the model name. There are two places to look, and they answer slightly different questions.
The Model Catalog in the console shows capability tags on every card — text to text, image text to text, browser use, computer use, speech to speech — alongside the deployment options (Serverless, Fine-Tuning, Self-Serve Deployment) and the context length and per-token prices. The left rail filters the grid: Type narrows to Text, Vision, or Voice models; Capabilities narrows by deployment option; Providers by vendor. Clicking a card opens the model page, where the Supported functionality table lists Image input, Function calling, context length, and max output explicitly.
The /v1/models endpoint returns the same information as JSON, which is what you want in code. Every model comes back with architecture.modality (text or multimodal), a tags list, a features list (for example tool_calling, structured_outputs, reasoning), and supported_parameters — the exact request fields it accepts. The reasoning feature flag matters for Step 3: it tells you which models think before they answer and therefore need a larger max_tokens.
curl -s "https://api.inference.crusoecloud.com/v1/models" \
-H "Authorization: Bearer $CRUSOE_API_KEY" | \
jq -r '.data[] | "\(.id)\t\(.architecture.modality)\t\(.tags | join(","))"'
Filter to models that accept images and tools:
curl -s "https://api.inference.crusoecloud.com/v1/models" \
-H "Authorization: Bearer $CRUSOE_API_KEY" | \
jq -r '.data[]
| select(.architecture.modality == "multimodal")
| select(.supported_parameters | index("tools"))
| .id'
⚠️ Warning: Model IDs are case-sensitive and the API expects the ID exactly as
/v1/modelsreturns it. Web pages do not always agree with the API: the docs listdeepseek-ai/DeepSeek-V4-Flash,zai/GLM-5.3-Flash, andqwen/Qwen3.8-27B, but the API only acceptsdeepseek-ai/Deepseek-V4-Flash(lowercase s),zai-org/GLM-5.3-Flash, andQwen/Qwen3.8-27B. Anything else returns404 model_not_found. Copy theidfield from the API response, not from a web page.
💡 Tip:
pricingin the same response has separateprompt,completion,input_cache_reads, andimagefields. At the time of writingimageis0for every vision model, meaning image input is billed as ordinary prompt tokens — but check it, because that is exactly the kind of thing that changes.
ℹ️ Note: Model names in this article are a snapshot and the catalog changes frequently. Treat
/v1/modelsand the Available Models docs page as the source of truth. For morejqfilters over the same response, see How-To List Model Features via curl and the Inference API.
Step 2: Configure the Client and Send Text
All examples share the same client. Text generation needs nothing beyond a messages array.
import os
import base64
import json
from openai import OpenAI
client = OpenAI(
base_url='https://api.inference.crusoecloud.com/v1',
api_key=os.getenv('CRUSOE_API_KEY'),
)
response = client.chat.completions.create(
model='deepseek-ai/Deepseek-V4-Flash',
messages=[
{'role': 'system', 'content': 'You are a concise assistant.'},
{'role': 'user', 'content': 'Summarize what MemoryAlloy does in two sentences.'},
],
)
print(response.choices[0].message.content)
Any chat model in the catalog works here. Pick by cost and quality — DeepSeek V4 Flash and Nemotron 3 Nano are the cheap high-volume options; Kimi K2.6 and DeepSeek V4 Pro are the heavy reasoners.
Step 3: Send an Image
Image input uses the OpenAI multi-part content format. content becomes a list of text and image_url parts instead of a string. The image_url can be a public HTTPS URL or a base64 data URI; the data URI route means the model never needs outbound network access to fetch your image.
VISION_MODEL = 'google/gemma-4-31b-it' # any image text to text model — see the Note below
def image_to_data_uri(path, mime='image/png'):
with open(path, 'rb') as f:
b64 = base64.b64encode(f.read()).decode()
return f'data:{mime};base64,{b64}'
response = client.chat.completions.create(
model=VISION_MODEL,
messages=[
{
'role': 'user',
'content': [
{'type': 'text', 'text': 'What GPU utilization does this Grafana panel show, and is anything wrong?'},
{'type': 'image_url', 'image_url': {'url': image_to_data_uri('dcgm_panel.png')}},
],
}
],
max_tokens=1024,
)
print(response.choices[0].message.content)
You can try the same thing without code in the Foundry chat playground — open any image text to text model, attach an image with the paperclip, and ask a question about it.
ℹ️ Note: Nemotron 3 Nano Omni, Kimi K2.6, Qwen3.8 27B, and GLM-5.3 Flash are reasoning models. They spend output tokens thinking before they answer, and with a small
max_tokens(under a few hundred) the budget runs out mid-thought andmessage.contentcomes backNone— which is easy to mistake for "the model ignored my image." Give reasoning models at leastmax_tokens=1024for image questions and checkfinish_reason:lengthmeans you starved it. The thinking itself comes back on the message as areasoningfield (Omni, Kimi) orreasoning_content(Qwen); read it withmessage.model_dump()since the OpenAI SDK does not type it. Gemma 4 answers directly with no reasoning phase and is the cheapest choice for quick captioning or extraction.
⚠️ Warning: Sending an image to a text-only model does not fail gracefully. Depending on the model you get a
400about unsupported content, or the model silently ignores the image and answers from the text alone — which looks like a working request returning nonsense. Confirmmodalityismultimodalin Step 1 first.
ℹ️ Note: Images count against context length as prompt tokens, and the cost scales with resolution — a small 400×200 image is on the order of 100 tokens, a 1280×800 screenshot on the order of 1,000. The exact count is reported per request in
usage.prompt_tokens_details. Downscale to the smallest resolution at which the content is still legible before encoding; 1024px on the long edge is a reasonable starting point for UI screenshots.
Step 4: Drive a Browser With Navigator n2
Navigator n2 works differently from the other models in this article, and getting the contract right matters more than the code.
Six things to know before you send a request:
-
You must pass the tool definition. On Crusoe,
yutori/n2is served as a standard OpenAI-compatible model; nothing is injected server-side. Send thecomputer_batchschema below intoolson every request. Without it, n2 answers in prose and describes actions it never took. -
It returns actions, not answers. A turn comes back as an assistant message with a
tool_callsentry namedcomputer_batch, whosearguments.actionsis an ordered list of{"name", "arguments"}primitives. When the task is done, n2 returns plaincontenttext and notool_calls— that is your loop's exit signal. -
Coordinates are normalized to a 1000×1000 space, origin top-left, regardless of what your schema says. The model is never told your resolution, so scale to real pixels before clicking:
px = x * width / 1000. -
Argument names are the model's, not yours. n2 chooses its own field names inside each action and is not always consistent: a click may arrive as
{"x": 172, "y": 333}or{"x": [547, 575]}; a key press as{"key": "ctrl+f"},{"command": ...}, or{"text": ...}; a scroll as{"direction": "down", "amount": 3}or{"scroll_down": 10}. Leave the per-actionargumentsschema open ({"type": "object"}), accept the variants in your executor, and log the raw arguments so you catch a new one. -
One tool result per tool call, carrying a screenshot taken after the last action in the batch. The endpoint accepts multipart
toolmessages (a list oftextandimage_urlparts), so the screenshot goes straight back in the tool result. -
At most two images per request. Send a third and the endpoint returns
400 At most 2 image(s) may be provided in one prompt. Before every call, replace theimage_urlparts in all but the two most recent image-bearing messages with a short text placeholder. Keep the text of every turn — the model needs its action history to avoid repeating work; it only needs the current and previous screen to act.
Standard OpenAI request fields (max_tokens, stream, tool_choice, response_format) all work with yutori/n2, the same as any other model in the catalog.
The description in the tool definition matters — it is how the model learns the coordinate convention and the action vocabulary you expect.
COMPUTER_BATCH_TOOL = {
'type': 'function',
'function': {
'name': 'computer_batch',
'description': (
'Run an ordered sequence of GUI actions on the desktop and return a screenshot '
'after the last one. Coordinates are [x, y] integers in a normalized 0-1000 space, '
'origin top-left. Execution stops at the first failing action.'
),
'parameters': {
'type': 'object',
'properties': {
'actions': {
'type': 'array',
'minItems': 1,
'items': {
'type': 'object',
'properties': {
'name': {'type': 'string', 'enum': [
'left_click', 'double_click', 'triple_click', 'middle_click', 'right_click',
'scroll', 'type', 'key_press', 'drag', 'mouse_move', 'wait', 'screenshot']},
'arguments': {'type': 'object'},
},
'required': ['name', 'arguments'],
},
},
},
'required': ['actions'],
},
},
}
The executor below maps those actions onto Playwright. It deliberately exposes no shell or filesystem tool — n2 was trained on a full desktop and will ask for one if offered, and a browser-only harness has nowhere safe to run it.
from playwright.sync_api import sync_playwright
N2_MODEL = 'yutori/n2'
W, H = 1280, 800 # viewport and screenshot must be the same size
KEYMAP = {'enter': 'Enter', 'backspace': 'Backspace', 'delete': 'Delete', 'tab': 'Tab',
'esc': 'Escape', 'space': 'Space', 'left': 'ArrowLeft', 'right': 'ArrowRight',
'up': 'ArrowUp', 'down': 'ArrowDown', 'pageup': 'PageUp', 'pagedown': 'PageDown',
'home': 'Home', 'end': 'End', 'ctrl': 'Control', 'alt': 'Alt', 'shift': 'Shift',
'meta': 'Meta', 'command': 'Meta', 'super': 'Meta'}
def to_px(args, prefix=''):
# Observed shapes: {"x": 172, "y": 333}, {"x": [547, 575]}, {"coordinates": [x, y]}. Accept all.
c = None
for k in ('coordinates', 'coordinate', 'position', 'point'):
if f'{prefix}{k}' in args:
c = args[f'{prefix}{k}']
break
if c is None:
x = args.get(f'{prefix}x')
c = x if isinstance(x, (list, tuple)) else [x, args.get(f'{prefix}y')]
if c is None or len(c) < 2 or c[0] is None or c[1] is None:
raise ValueError(f'no coordinates in {args}')
x, y = float(c[0]), float(c[1])
return min(max(x, 0), 1000) * W / 1000, min(max(y, 0), 1000) * H / 1000
def get_keys(args):
# Observed shapes: {"key": ...}, {"command": ...}, {"text": ...}.
for k in ('key', 'keys', 'command', 'text', 'combo'):
if args.get(k):
return args[k]
raise ValueError(f'no key in {args}')
def scroll_delta(args):
# Observed shapes: {"direction": "down", "amount": 3} and {"scroll_down": 10}.
if 'scroll_down' in args:
return args['scroll_down'] * 100
if 'scroll_up' in args:
return -args['scroll_up'] * 100
amount = args.get('amount', 3)
return amount * 100 * (-1 if args.get('direction') == 'up' else 1)
def pw_key(k):
return '+'.join(KEYMAP.get(part, part) for part in k.split('+'))
def run_action(page, name, args):
if name in ('left_click', 'double_click', 'triple_click', 'right_click', 'middle_click'):
x, y = to_px(args)
clicks = {'double_click': 2, 'triple_click': 3}.get(name, 1)
button = {'right_click': 'right', 'middle_click': 'middle'}.get(name, 'left')
page.mouse.click(x, y, button=button, click_count=clicks)
elif name == 'type':
page.keyboard.type(args['text'])
elif name == 'key_press':
for combo in get_keys(args).split(' '):
page.keyboard.press(pw_key(combo))
elif name == 'scroll':
x, y = to_px(args)
page.mouse.move(x, y)
page.mouse.wheel(0, scroll_delta(args))
elif name == 'drag':
sx, sy = to_px(args, prefix='start_'); ex, ey = to_px(args)
page.mouse.move(sx, sy); page.mouse.down(); page.mouse.move(ex, ey); page.mouse.up()
elif name == 'mouse_move':
page.mouse.move(*to_px(args))
elif name == 'wait':
page.wait_for_timeout(args.get('duration', 1) * 1000)
elif name == 'screenshot':
pass # every batch returns a screenshot anyway
else:
raise ValueError(f'unsupported action: {name}')
def screenshot_uri(page):
page.screenshot(path='shot.jpg', type='jpeg', quality=80)
return image_to_data_uri('shot.jpg', mime='image/jpeg')
def trim_images(messages, keep=2):
# The endpoint rejects requests with more than 2 images. Keep the two most recent screenshots;
# replace older ones with a text placeholder so the action history survives intact.
seen = 0
for m in reversed(messages):
content = m.get('content')
if isinstance(content, list) and any(p.get('type') == 'image_url' for p in content):
seen += 1
if seen > keep:
m['content'] = [p if p.get('type') != 'image_url'
else {'type': 'text', 'text': '[earlier screenshot removed]'}
for p in content]
HARNESS_NOTE = (
'Environment: a single Chromium browser tab. There is no desktop, terminal, or other application; '
'keyboard shortcuts that open system tools will fail. Use only clicks, typing, scrolling, and '
'in-page keys. When the task is complete, reply with the answer as plain text and no tool call.'
)
def run_n2(task, start_url, max_turns=30):
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
page = browser.new_page(viewport={'width': W, 'height': H})
page.goto(start_url)
messages = [{'role': 'user', 'content': [
{'type': 'text', 'text': f'{task}\n\n{HARNESS_NOTE}'},
{'type': 'image_url', 'image_url': {'url': screenshot_uri(page)}},
]}]
for turn in range(max_turns):
trim_images(messages, keep=2) # endpoint limit: at most 2 images per request
resp = client.chat.completions.create(
model=N2_MODEL,
messages=messages,
tools=[COMPUTER_BATCH_TOOL],
max_tokens=4096,
)
msg = resp.choices[0].message
messages.append(msg.model_dump(exclude_none=True))
if not msg.tool_calls:
browser.close()
return msg.content # task complete
for call in msg.tool_calls:
args = json.loads(call.function.arguments)
if call.function.name != 'computer_batch':
messages.append({'role': 'tool', 'tool_call_id': call.id,
'content': f'ERROR: {call.function.name} is not available in this harness.'})
continue
actions = args.get('actions', [])
done, error = 0, None
for a in actions:
try:
run_action(page, a['name'], a.get('arguments', {}))
done += 1
except Exception as e:
error = f'{a["name"]} failed: {e}'
break # stop at first error, do not run the rest
page.wait_for_timeout(500)
text = f'Executed {done} of {len(actions)} actions.' + (f' {error}' if error else '')
messages.append({'role': 'tool', 'tool_call_id': call.id, 'content': [
{'type': 'text', 'text': text},
{'type': 'image_url', 'image_url': {'url': screenshot_uri(page)}},
]})
# Out of turns: ask for a summary instead of dropping the run mid-trajectory.
messages.append({'role': 'user', 'content': 'Stop and summarize your progress so far.'})
trim_images(messages, keep=2)
final = client.chat.completions.create(model=N2_MODEL, messages=messages, max_tokens=1024)
browser.close()
return final.choices[0].message.content
print(run_n2(
task='Find the context length listed for Kimi K2.6 on this page and report it.',
start_url='https://docs.crusoecloud.com/serverless-inference/available-models',
))
⚠️ Warning: Model-generated actions are untrusted input. Never point this harness at a browser session that is logged into anything you care about, clamp coordinates to the viewport (the
to_pxhelper does), and if you do add abashtool, run it in a throwaway container. A confused model will happily click "Delete".
ℹ️ Note:
playwright install chromiumdownloads a browser fromcdn.playwright.devand times out behind some corporate proxies. If that happens, skip the download and drive the Chrome you already have:p.chromium.launch(headless=True, channel='chrome'). Everything else in the code is unchanged.
💡 Tip:
trim_imagesand the prefix cache pull in opposite directions. The cache serves the unchanged prefix of a request; the trimmer rewrites an earlier message every time it drops a screenshot, and everything after that point stops matching. Expectcached_tokensto climb for a couple of turns, reset when an old image is dropped, and rebuild. You cannot avoid the trim (two-image limit), and at roughly 1,000 tokens per screenshot it is also what keeps a long run inside the context window — so treat the cache as a discount you get most turns, not a guarantee.
ℹ️ Note: n2 was trained on a full desktop, and it shows: it will reach for
ctrl+fto find text orctrl+alt+tto open a terminal, neither of which exists in a headless browser tab. TheHARNESS_NOTEappended to the task tells it there is no desktop; when it tries anyway, the executor reports the failure in the tool result and the model adapts on the next turn — typically by scrolling. Do not silently swallow unsupported actions — the error text is how it learns the boundaries of your harness.
ℹ️ Note: You can add your own tool definitions alongside
computer_batch— alookup_orderorpost_to_slackfunction, say — and n2 will call them the same way, since on this endpoint it is ordinary OpenAI-style function calling.
Step 5: Drive a Browser With a General Vision Model (Alternative)
If you need a model n2 isn't — an open-weights one you can also fine-tune, or one that doubles as your general assistant — the same loop works with any multimodal model that lists tools in supported_parameters. You define the action vocabulary yourself and the model returns ordinary OpenAI tool_calls.
The code below reuses client and image_to_data_uri from Steps 2–3 and sync_playwright, W, H, and screenshot_uri from Step 4, so run it in the same file.
BROWSER_TOOLS = [
{'type': 'function', 'function': {
'name': 'click',
'description': f'Click at PIXEL coordinates in the {W}x{H} screenshot. x in 0-{W}, y in 0-{H}.',
'parameters': {'type': 'object',
'properties': {'x': {'type': 'integer'}, 'y': {'type': 'integer'}},
'required': ['x', 'y']}}},
{'type': 'function', 'function': {
'name': 'type_text', 'description': 'Type text into the focused element.',
'parameters': {'type': 'object', 'properties': {'text': {'type': 'string'}}, 'required': ['text']}}},
{'type': 'function', 'function': {
'name': 'done', 'description': 'Call when the task is complete, with the final answer.',
'parameters': {'type': 'object', 'properties': {'answer': {'type': 'string'}}, 'required': ['answer']}}},
]
SYSTEM = (f'You control a web browser. Each turn you are shown a {W}x{H} pixel screenshot. '
'Respond ONLY with a tool call using pixel coordinates from the screenshot. '
'Call done() when finished.')
# Observed coordinate conventions against the pixel schema above. Calibrate before trusting.
COORD_MODE = {
'google/gemma-4-31b-it': 'norm1000', # returned 0-1000 despite pixel schema
'zai-org/GLM-5.3-Flash': 'pixels', # returned true pixels
'nvidia/Nemotron-3-Nano-Omni-Reasoning-30B-A3B': 'pixels', # pixels when the size is stated; 0-1 floats when not
}
def to_pixels(x, y, mode):
if 0 <= x <= 1 and 0 <= y <= 1: # 0-1 floats are unambiguous at any viewport
return x * W, y * H
if mode == 'norm1000':
return x * W / 1000, y * H / 1000
return x, y
def calibrate(model):
"""One-shot check: ask for a click on a known target and infer the convention from the answer."""
from PIL import Image, ImageDraw
img = Image.new('RGB', (W, H), 'white')
ImageDraw.Draw(img).rectangle([W * 0.7, H * 0.7, W * 0.8, H * 0.8], fill='black') # centre ≈ (0.75W, 0.75H)
img.save('calib.png')
resp = client.chat.completions.create(
model=model, tools=BROWSER_TOOLS, tool_choice='auto', max_tokens=1024,
messages=[{'role': 'system', 'content': SYSTEM},
{'role': 'user', 'content': [
{'type': 'text', 'text': 'Click the black square.'},
{'type': 'image_url', 'image_url': {'url': image_to_data_uri('calib.png')}}]}],
)
tc = resp.choices[0].message.tool_calls
if not tc:
raise RuntimeError(f'{model} returned no tool call during calibration')
a = json.loads(tc[0].function.arguments)
x, y = a['x'], a['y']
if 0 <= x <= 1 and 0 <= y <= 1:
return 'pixels' # floats are handled directly by to_pixels
return 'pixels' if abs(x - 0.75 * W) < 0.1 * W else 'norm1000'
def run_generic(task, start_url, model=VISION_MODEL, max_steps=25):
mode = COORD_MODE.get(model) or calibrate(model)
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
page = browser.new_page(viewport={'width': W, 'height': H})
page.goto(start_url)
messages = [{'role': 'system', 'content': SYSTEM}]
for step in range(max_steps):
messages.append({'role': 'user', 'content': [
{'type': 'text', 'text': f'Task: {task}\nStep {step + 1}.'},
{'type': 'image_url', 'image_url': {'url': screenshot_uri(page)}},
]})
resp = client.chat.completions.create(
model=model, messages=messages, tools=BROWSER_TOOLS, tool_choice='auto',
temperature=0.1, max_tokens=1024,
)
msg = resp.choices[0].message
messages.append(msg.model_dump(exclude_none=True))
if not msg.tool_calls:
browser.close(); return msg.content
for call in msg.tool_calls:
args = json.loads(call.function.arguments)
if call.function.name == 'click':
page.mouse.click(*to_pixels(args['x'], args['y'], mode))
elif call.function.name == 'type_text':
page.keyboard.type(args['text'])
elif call.function.name == 'done':
browser.close(); return args['answer']
page.wait_for_timeout(500)
messages.append({'role': 'tool', 'tool_call_id': call.id, 'content': json.dumps({'ok': True})})
browser.close()
return 'Step limit reached without done().'
Expect to spend time here that n2 makes unnecessary.
Coordinate conventions are the main problem: given the same pixel-integer schema, one model returns pixels, another returns 0–1000 normalized integers, another returns 0–1 floats — and a model may switch depending on whether you told it the image size. You cannot tell 0–1000 values from pixels by magnitude once your viewport is wider than 1000px, which is why the code carries a per-model COORD_MODE and a one-shot calibrate() that clicks a known target before the first real task, and why the pixel dimensions are stated in both the tool description and the system prompt.
Forcing a specific function with tool_choice={'type': 'function', 'function': {'name': 'click'}} makes the model call it but does not make the arguments sensible; validate every argument before acting on it.
You also own the history-trimming logic: drop image_url parts from turns more than a few steps old or context will run out around step 15.
Common Issues
-
404with"code": "model_not_found"— ID case or org prefix is wrong. Copy theidfrom/v1/models, not from a web page; see the Warning in Step 1 for the known mismatches. -
400mentioningcontenttype,image_url, or unsupported input — the model is text-only. Re-run Step 1 and pick amultimodalmodel. -
message.contentisNoneor empty after an image question — a reasoning model ran out ofmax_tokenswhile thinking. Checkfinish_reason == 'length'and raisemax_tokensto 1024 or more. -
yutori/n2replies in prose claiming it clicked something — you did not passtools=[COMPUTER_BATCH_TOOL]. The model has nothing to call, so it narrates. - n2 clicks land in the wrong place — you forgot to denormalize from the 1000×1000 space, or the viewport and screenshot dimensions differ.
W/Hmust match what you pass tonew_page(viewport=...). - n2 repeats the same action every turn — the tool result's screenshot is stale (taken before the batch ran) or you dropped earlier turns from
messages. Send the text of every turn; screenshot after the last action. -
400 At most 2 image(s) may be provided in one prompt— you sent three or more screenshots in one request. Calltrim_imagesbefore every request; it keeps the two most recent and replaces the rest with text. -
KeyErroror'NoneType' object has no attribute 'split'in your n2 executor — the model used an argument name your handler does not know. Logcall.function.argumentson every action and extendto_px,get_keys, orscroll_deltawhen a new one appears. - n2 presses
ctrl+alt+t,alt+tab, or other desktop shortcuts — it thinks it has a desktop. Append theHARNESS_NOTEto the task and return the failure in the tool result; do not swallow it. -
400abouttoolmessage ordering — the assistant turn containingtool_callswas dropped before thetoolresult was appended. Always appendmsgfirst. - General VLM describes the screenshot but never emits a
tool_call— it may not supporttools(checksupported_parametersin Step 1), or it is a reasoning model with too small amax_tokens. Raise the budget before reaching for forcedtool_choice. - Clicks from a general VLM land near the top-left corner, or consistently short of the target — the model used a different coordinate convention from your schema. Run
calibrate()for that model, setCOORD_MODEfrom the result, and logcall.function.argumentsif you see a convention other than pixels, 0–1 floats, or 0–1000 integers. -
429 Too Many Requests— agent loops multiply request volume. See How-To Fix 429 RateLimitError When Interacting With Crusoe Managed Inference Service.
Example
A platform team runs a nightly check that verifies their public status page and pricing page render correctly after deploys. Before, it was a brittle Selenium script with hard-coded CSS selectors that broke on every frontend change.
They replace it with the Step 4 harness running yutori/n2 on Crusoe Managed Inference. Each night the agent is handed a task like "Open the pricing page, confirm the H200 hourly price is displayed, and report it." n2 looks at the screenshot, scrolls or clicks as needed, and returns the price as its final text. Because it reasons over pixels instead of selectors, a redesign that would have broken the old script becomes a slightly different screenshot.
For a sense of scale: the Step 4 code, pointed at the Crusoe docs "Available models" page and asked for Kimi K2.6's context length, tried ctrl+f, noticed no find bar appeared, scrolled twice, and answered 256K — four turns, about twelve seconds.
The same team uses text-only deepseek-ai/Deepseek-V4-Flash for the post-processing step — turning the night's answers into a Slack summary — because that step never needs to see an image and Flash is a fraction of the price.
Related Articles
- How-To Get Started With Text Generation on Crusoe Managed Inference (Python)
- How-To List Model Features via curl and the Inference API
- How-To Test the Prefix Cache Hit Rate on Serverless Inference
- How-To Fix 429 RateLimitError When Interacting With Crusoe Managed Inference Service
- How-To Hook Up OpenClaw to Crusoe Managed Inference