Ocean Network Inference

Models, apps and templates, on a card you book by the hour

Point us at any OpenAI compatible Hugging Face model, open an app and set it up yourself, or start from a template that is already configured. The GPU is yours alone while it runs.

H200 from $2.16 an hour
OpenAI-compatible endpoint
Billed for the window you book
Scroll
Most people start with one of these, though anything on the Hub works
Most people start with one of these, though anything on the Hub works
Qwen3-8BQwen3-32BQwen3-Coder-30Bgpt-oss-120bDeepSeek-V4-FlashLlamaMistralGemmaPhi-4Muse-Glimmer-30BMiniMax H3LTX VideoStable DiffusionNomic Embedyour own repoQwen3-8BQwen3-32BQwen3-Coder-30Bgpt-oss-120bDeepSeek-V4-FlashLlamaMistralGemmaPhi-4Muse-Glimmer-30BMiniMax H3LTX VideoStable DiffusionNomic Embedyour own repo
Models

Start with a model, either one we have already sized for the hardware

Curated packages pair a model with an engine preset already checked against the card it asks for, so nothing dies at startup. Anything else on the Hub goes through the custom flow, and both routes answer on the OpenAI API your code already calls.

text generation
$ curl $ONC/v1/chat/completions
-d '{"model":"Qwen3.8-27B","stream":true}'
live · H200 · stops at the end of your window
Services

Or bare apps that you can configure yourself

Services are the bare apps. You get an empty ComfyUI, OpenWebUI, OpenCode or Hermes with nothing configured in it, then you go to Models, bring up the one you actually want, and set it up inside the app yourself.

comfyui · port 8188
checkpoint loaded · sampling 28 steps
Templates

Or a template that arrives with the weights already downloaded

Templates are those same apps with the setup already finished. We have chosen the models and built the workflow, so it runs the moment it starts.

template · image to video
weights already on the bucket
Comparison

Why people move their inference to Ocean Network

The questions that come up every time somebody weighs us against what they use today.

What you getOcean NetworkServerless APIsManaged endpointsRaw GPU rental
The model you name is the model that loadsYesPooled, provider not pinnedYesYes, if you build it
A card nobody else is sharingYesNoYesYes
Price agreed before anything startsHourly, shown upfrontPer token, after the factHourlyHourly
Weights kept between runsOptional bucketNot applicableRe-pulled on cold startIf you wire it up
Running without configuring anythingTemplates, services and curated models, one clickYesSome setupYou build the image
Pricing

H200 snapshot, per hour, per card

You pay a node operator directly for the hardware and the minutes you use, and the window you book is a hard stop.

Ocean Networkread 19 Aug 2026
$2.16
RunPod Secure Cloudread 26 Aug 2026
$4.39
RunPod H200 SXM141 GB VRAM, 276 GB RAM, 24 vCPU
$4.59
Hugging Face Endpointsread 30 Jul 2026
$5.00
Together dedicated clusterread 30 Jul 2026
$5.99
Specifications differ between providers, so read the card and the memory alongside the rate. Payment here is a deposit and an authorisation rather than a transfer: the node claims against it as your service runs.
What the market was missing

You told us what’s broken with legacy providers. So we built the solution.

“If I cannot pin the model provider, I do not want it.”
Top reply, 188 points, under a thread about the same model returning four different qualities from four providers
You can pin it here. The model you name is the model that loads, on hardware nobody else is touching, and the hourly number is on the screen before you press start.

If that sounds like what you have been looking for

Pick a package, an app, or any repository on the Hub, and pay for the window you booked. The docs walk through all three routes if you would rather read first.