

Scaleway
on Orq.ai
Use Scaleway’s Generative API and managed inference through a single Orq.ai integration. Route open‑source models such as Llama 3.3, Mistral Medium/Small, Qwen, Gemma, Pixtral, and audio/embedding models via Orq’s AI Router for chat, reasoning, coding, vision, and GPU‑backed workloads.
Capabilities:
Chat
Reasoning
Speech
Vision
Models Supported:
glm-5.2
gemma-4-26b-a4b-it
qwen3.6-35b-a3b
devstral-2-123b-instruct-2512
gpt-oss-120b
Provider HQ:
Scaleway, a European cloud provider headquartered in France.
Access Scaleway through Orq.ai’s AI Router
Scaleway is a European cloud provider offering two main AI surfaces: a Generative API (serverless model‑as‑a‑service billed per 1M tokens) and Managed Inference on dedicated GPUs billed hourly. These endpoints serve popular open‑source families (Llama, Mistral, Qwen, Gemma, etc.) behind an OpenAI‑compatible API with EU‑based infrastructure.
Orq.ai supports Scaleway’s Generative API via an OpenAI‑compatible provider configuration, with availability depending on your Scaleway account, region, and workspace setup.
Scaleway models available on Orq.ai
Model (example) | Type | Context / scope (tokens) | Best for | Pricing tier (reference, Paris Generative API) |
|---|---|---|---|---|
mistral-medium‑3.5‑128b | Chat / reasoning / vision | Up to around 256K context | Complex reasoning, long‑context chat, and high‑stakes tasks on a strong Mistral flagship model | Premium – €1.50 / 1M input tokens and €7.50 / 1M output tokens, 50% discount via Batches |
qwen3.5‑397b‑a17b, qwen3.6‑35b‑a3b | Chat / code / vision | Up to 262K context (Qwen3.x families) | Everyday production workloads, multilingual chat, RAG, and coding where Qwen offers strong price/performance | Mid‑tier – examples: qwen3.5‑397b at €0.60 / 1M input and €3.60 / 1M output; qwen3.6‑35b at €0.25 / 1M input and €1.50 / 1M output |
gpt‑oss‑120b, mistral‑small‑3.2‑24b, pixtral‑12b | Chat / fast / cost‑efficient / vision | 128K–131K context (depending on model) | Fast, lower‑cost tasks such as everyday chat, classification, extraction, lightweight vision, and high‑volume support | Cost‑efficient – examples: gpt‑oss‑120b at €0.15 / 1M input and €0.60 / 1M output; mistral‑small‑3.2‑24b at €0.15 / 1M input and €0.35 / 1M output; pixtral‑12b at €0.20 / 1M input and €0.20 / 1M output |
Why use Scaleway through Orq.ai?
Capability | Provider / models | Direct | Through Orq.ai |
|---|---|---|---|
Chat & reasoning | Llama 3.3, Mistral Medium/Small, Qwen3.x, Gemma 3/4, GPT‑OSS | Call Scaleway’s Generative API directly via its OpenAI‑compatible endpoints for chat and reasoning | Use Scaleway models through Orq.ai’s router, adding central routing, tracing, evals, budgets, and governance around each request |
Coding | qwen3‑coder‑30b, devstral‑2‑123b, GPT‑OSS‑120B | Use Scaleway directly for code generation, debugging, and refactoring on open‑source code‑tuned models | Route coding workloads through Orq.ai, compare Scaleway‑served models against other providers, and monitor cost, latency, and quality from one control layer |
Vision / audio / embeddings | Pixtral‑12b, Gemma Vision models, Whisper, Qwen embeddings | Use Scaleway’s Generative API for vision and audio, and embeddings directly in your apps | Use Orq.ai to orchestrate when to call Scaleway (for example, EU‑hosted vision or embeddings) while other steps use alternative providers, all under unified observability and control |
This gives teams a practical way to use Scaleway where it performs best while centralising routing, observability, evals, and cost controls across the wider model stack.
Pricing
Scaleway offers:
Generative API: token‑based pricing per 1M tokens for chat, code, vision, audio, and embeddings, with a free tier
Managed Inference: dedicated GPU instances billed hourly for custom model deployment
For Generative API (Paris region; 2024–2026 data):
Free tier: the first 1,000,000 tokens and about 60 minutes of audio transcription are free; billing starts from token 1,000,001
Batches API: requests sent via Batches receive a 50% discount on token rates
Representative token prices (input / output per 1M tokens):
mistral‑medium‑3.5‑128b: €1.50 in / €7.50 out
llama‑3.3‑70b‑instruct: €0.90 in / €0.90 out
qwen3.5‑397b‑a17b: €0.60 in / €3.60 out
gpt‑oss‑120b: €0.15 in / €0.60 out
mistral‑small‑3.2‑24b‑instruct‑2506: €0.15 in / €0.35 out
devstral‑2‑123b‑instruct‑2512: €0.40 in / €2.00 out
pixtral‑12b‑2409: €0.20 in / €0.20 out
qwen3‑embedding‑8b / bge‑multilingual‑gemma2: €0.10 / 1M embedding tokens, output free
Managed Inference hourly GPU pricing (Paris, examples):
L4‑1‑24G: about €0.93/hour.
L40S‑1‑48G: about €1.72/hour.
H100‑1‑80G: about €3.40/hour.
Multi‑GPU configs (2–8× H100 / H100‑SXM) scale linearly, with hourly costs from roughly €6.68 to €30.06/hour depending on configuration
Check the Orq.ai pricing page and your workspace’s provider configuration for current per‑model and per‑GPU rates when using Scaleway via Orq.ai.
Compatible frameworks and tools
Orq.ai exposes Scaleway through:
a Scaleway provider configuration in AI Gateway > BYOK, where you paste your Scaleway API key, and
an OpenAI‑compatible integration, since Scaleway’s Generative API already uses that format
That means:
Existing backend workflows using OpenAI SDKs or OpenAI‑compatible clients can switch to Orq.ai’s base URL and route some traffic to Scaleway‑hosted open‑source models without changing request formats.
Agents, code assistants, and tools that integrate with Orq.ai can be configured so that EU‑hosted or open‑source steps are served by Scaleway, while other steps use different providers, all sharing the same observability and governance layer.
Check the Orq.ai integration docs for the latest supported frameworks and tools for Scaleway.
FAQs
Do I need a separate Scaleway account to use Scaleway through Orq.ai?
Yes. You create a Scaleway API key in the Scaleway console, then configure it in Orq.ai’s AI Gateway; in some cases, usage billed via Orq.ai may also be available depending on plan and region. In both setups, Orq.ai provides one place to manage routing, observability, and cost controls around that Scaleway usage.
Can I route only some workflows to Scaleway and others to different providers?
Yes. You define routes per workflow in Orq.ai and choose which ones should use Scaleway vs other providers, so you can reserve Scaleway for EU data‑residency, open‑source, or specific price/performance needs while sending other tasks elsewhere.
Does using Scaleway through Orq.ai add latency?
Orq.ai is designed as a lightweight router, so the added overhead is small compared to model inference time. You can use routing policies, caching, and provider selection to keep end‑to‑end performance within your targets while gaining visibility and control.
Alternatives to
Scaleway
Anthropic
Use Claude Opus 4.7, Sonnet 4.6, Haiku 4.5, and other supported Claude models through one API.
Chat
Code
Reasoning
Vision
Models:
claude-opus-5
claude-sonnet-5
claude-fable-5
Open AI
Use OpenAI's foundation models through a single Orq.ai API. Route models such as GPT-4.1, GPT-4.1-mini, o3-mini, and GPT-4o-class models via Orq's AI Router for chat, reasoning, coding, and multimodal workloads.
Chat
Code
Image Generation
Reasoning
Speech
Vision
Models:
gpt-5.5 (EU)
gpt-5.6-luna (EU)
gpt-5.6-sol (EU)
Google AI
Use Google’s Gemini models through a single Orq.ai API. Route models such as Gemini 3.1 Pro, Gemini 2.5 Flash, and Gemini 2.0 Flash‑Lite via Orq’s AI Router for chat, reasoning, coding, and multimodal workloads.
Chat
Code
Embeddings
Image Generation
Reasoning
Speech
Vision
Models:
gemini-3.5-flash-lite (Gemini API)
gemini-3.6-flash (Gemini API)
Gemini 3 Pro Image
AWS
Use AWS Bedrock’s foundation models through a single Orq.ai API. Route models such as Amazon Nova, Amazon Titan Text, and compatible third‑party models exposed via Bedrock through Orq’s AI Router for chat, reasoning, coding, and multimodal workloads.
Chat
Code
Embeddings
Reasoning
Vision
Models:
eu.anthropic.claude-opus-5
global.anthropic.claude-opus-5
us.anthropic.claude-opus-5


