Skip to content

The layer
above the card

Notebook, training run, endpoint and finished models — as a service instead of an empty machine. A training run trains a model on your data. An endpoint is the fixed network address where a finished model responds. The underlying hardware is the same that is rented by the hour in the GPU catalogue — you work on the model, not on the machine.

Prices are net plus 19% VAT., location Frankfurt. If you want the card without the platform on top, you can find it under GPU.

What the platform delivers
Accelerator
€0.7397$0.8581 to €5.2500$6.0900 per card and hour
Workspaces
Jupyter and VS Code
Training runs
1 to 8 nodes
Endpoints
Scale to zero
Model catalogue
6 models, open weights
Billing
per second or per token
Inference locations
4 in Germany and Finland
All amounts on this page are derived from the hourly rates of the NG instance classes. The calculation is shown below each table.
Accelerators from
€0.7397$0.8581
per card and hour, L4 with 24 GB
Workspace from
€0.0567$0.0658
per hour, billed by the second
Output from
€0.1251$0.1451
per million tokens — a token is roughly a part of a word
Inference locations
4
Inference means: a finished model responds. No processing outside the EU.

Three layers, three billing units

Accelerator, service and model are not the same product. They differ in what you need to bring and how they are billed.

  1. Layer 1 · Hardware

    Rent accelerators

    Billing per hour, per card

    You get an instance with one, two, four or eight cards, root access and an empty machine with drivers. Operating system, framework, model state and scaling are up to you. You can find the hourly rate further down this page.

    Right for you if you have your own toolchain and need full control over the environment.

    View GPU instances (separate page)

  2. Layer 2 · Platform

    Notebook, training run, endpoint

    Billing per service runtime

    The same hardware, but as a service: Jupyter starts ready to use, a training run is queued instead of logged in, a model becomes an HTTPS endpoint with scale-to-zero.

    Right for you if you work on models and not on machines.

    To the workspaces

  3. Layer 3 · Model

    Call ready-to-use models

    Billing per million tokens

    No node, no deployment, no waiting time. A call against a model name; you are billed for what went through the wire.

    Right for you if you are building an application and do not want to operate a model.

    To the model catalogue

What an hour of computing costs

Five accelerator classes support every row on this page where a card appears — from the accelerated workspace to the endpoint. The amounts are taken unchanged from the instance catalogue. Only the reference unit is new here: the rate per node is broken down to the individual card and to the hour.

What a card can support is determined by its graphics memory — the VRAM, the memory on the card itself. It ranges from 24 GB on the L4 to 141 GB on the H200 SXM. A model whose weights do not fit will not run on this card — which is why the memory size is in the same row as the price.

Four accelerator cards vertically in a dense compute node, milled cooling fins, bundled power cables from the right.
A L4 with 24 GB costs €0.7397$0.8581 per hour at ENTRONYX CLOUD, a H100 SXM with 80 GB costs €2.6575$3.0827. Billing is per second at the hourly rate, capped at the monthly price — you will never be charged for more than 730 hours a month.
Graphics memory, hourly rate and monthly rate of the five accelerator classes, per card and per node
ClassGraphics memoryCards per nodeCard / hNode / hNode / month
NG-L4Inference of small and quantised models, embeddings24 GB per card1€0.7397$0.8581€0.7397$0.8581€540.00$626.40
NG-L40SModels up to around 32 billion parameters, image models48 GB per card1€1.3808$1.6017€1.3808$1.6017€1,008.00$1,169.28
NG-RTX6000Mixture of experts models and long contexts — 96 GB VRAM96 GB per card1€1.6425$1.9053€1.6425$1.9053€1,199.00$1,390.84
NG-H100Training runs — the reference class for all calculations below80 GB per card1€2.6575$3.0827€2.6575$3.0827€1,940.00$2,250.40
NG-H200×8Large models across multiple cards, distributed training141 GB per card8only as a whole node€5.2500$6.0900€42.0000$48.7200€30,660.00$35,565.60

All amounts net, plus 19% VAT., Frankfurt location. The monthly rate is maintained; the hourly rate is this amount divided by 730 hours. The four single-card classes can be ordered individually; the eight-card class is only allocated as a whole node, so its card rate is a calculated value and not a separate quote.

The platform either adds a named surcharge to these rates — for the workspaces — or converts them into a token price via a measured throughput. A token is the building block into which a model breaks down text — usually a short word or a syllable. Both calculation methods are shown below the tables in the following sections.

A notebook that is already running

A workspace is an instance with a ready-to-use interface: Jupyter and VS Code are set up, the drivers are correct, and Object Storage is mounted as a directory. This is where fine-tuning happens — retraining a pre-trained model on your own data. You are billed for runtime, not ownership.

Jupyter and VS Code
Both interfaces run on the same environment and can be switched during operation. VS Code is also accessible via the desktop editor's remote connection.
Accelerators can be added
The environment starts without a GPU and can be upgraded to an accelerated class at the next start. The working directory is retained, the price changes with the class.
Billing by runtime
The time the environment is running is counted — to the second, not the time it is created. A stopped environment only costs its storage.
Data connection to Object Storage
Buckets are mounted as a directory at startup. Access keys are stored in the project, not in the notebook; they do not appear in the cell history or in checkpoints.

What a workspace costs

The hourly rate of a workspace is the hourly rate of its instance class plus a 15% platform fee. This fee covers image maintenance, operating both interfaces, idle detection, and the S3 connection, i.e. access to Object Storage. The instance price itself remains the same as in the catalogue — the platform does not make the card more expensive, it adds a specified amount alongside it.

Hourly rate and monthly cap of workspaces by instance class, split into instance price and platform fee
ClassConfigurationAcceleratorInstance / hPlatform / hEnvironment / hCap / month
NP3Data preparation, analysis, classic machine learning without accelerators4 vCPU · 8 GB RAMwithout accelerator€0.0493$0.0572€0.0074$0.0086€0.0567$0.0658€41.39$48.01
ND3The same work with constant performance — no shared CPU time4 vCPU · 16 GB RAMwithout accelerator€0.1063$0.1233€0.0159$0.0184€0.1222$0.1418€89.21$103.48
NG-L4First experiments with quantised language models, embeddings, image models8 vCPU · 48 GB RAM1 × L4€0.7397$0.8581€0.1110$0.1288€0.8507$0.9868€621.01$720.37
NG-L40SLoRA adaptations and models up to around 30 billion parameters16 vCPU · 96 GB RAM1 × L40S€1.3808$1.6017€0.2071$0.2402€1.5879$1.8420€1,159.17$1,344.64
NG-H100Full fine-tuning of medium-sized models, HPC workloads with FP6432 vCPU · 256 GB RAM1 × H100 SXM€2.6575$3.0827€0.3986$0.4624€3.0561$3.5451€2,230.95$2,587.90

All amounts net, plus 19% VAT., Frankfurt location. The monthly cap corresponds to 730 hours: a continuously running workspace never costs more, regardless of the number of calendar days.

Four accelerator cards in an open node, the fans spinning slowly behind their grilles.

A run is billed by runtime at ENTRONYX CLOUD, not by result — even a cancelled one costs the seconds it ran. However, it resumes from the last complete checkpoint and not from the beginning.

A run is a job, not a session

You do not log in to a machine, but submit a job. The platform finds the nodes, mounts the data, writes checkpoints — secured intermediate states from which a run can resume — and releases the cards again as soon as the last step has run.

5 steps

Job flow

  1. 01SubmitA job consists of an image, command, data source and the desired class. It is submitted via the console, the API or from a pipeline — a session is not required for this.
  2. 02QueueIf the class is occupied at the selected location, the job waits instead of failing. The estimated start time is shown in the job; jobs in a project queue run in order.
  3. 03ExecuteAccelerators, storage and network are allocated, the image is started, the data source is mounted. Across multiple nodes, the platform sets the rank and the addresses of the peers itself.
  4. 04CheckpointsA checkpoint directory is written to Object Storage. An aborted job resumes at the last complete checkpoint, not at step zero.
  5. 05Log and endStandard output, metrics and events remain accessible for 30 days and can be downloaded as a file. With the last step, the nodes are released and billing ends.

Which model size runs on how many cards

The weights of a model must fit into the graphics memory. At 8 bits — quantisation, i.e. shortening each number to one byte — one billion parameters occupy around one gigabyte; at BF16 with two bytes per number, it is double. Therefore, a card with 24 GB is sufficient for a model with eight billion parameters, while a model with seventy billion needs two cards. The rows below are the allocation with which the model catalogue on this page actually runs.

Number of parameters, size of weights, number of cards per replica and available graphics memory per model in the catalogue
ModelParametersNumber formatWeightsCards per replicaGraphics memory
Ministral 8B Instruct8 billion8 bit8 GB1 × L424 GB
Mistral Small Instruct24 billion8 bit24 GB1 × L40S48 GB
Qwen2.5 Coder 32B32 billion8 bit32 GB1 × L40S48 GB
Mixtral 8×7B Instruct47 billion, 13 billion active8 bit47 GB1 × RTX PRO 6000 Blackwell Max-Q96 GB
Llama 3.3 70B Instruct70 billion8 bit70 GB2 × H200 SXM282 GB
BGE-M3568 millionBF161.2 GB1 × L424 GB

A replica is a running copy of the model. The graphics memory must hold the cache of running requests in addition to the weights — which is why the column on the right is consistently larger than the weights column and not tightly sized.

What distribution costs

The reference is a run over 72 hours on a node of class NG-H100 at €2.6575$3.0827 per hour — the same calculation as on the GPU page, just distributed across multiple nodes here. More nodes shorten the wait time, but they do not lower the price: gradient exchange — the synchronisation of learning steps between nodes — costs efficiency, and you pay for GPU hours.

Elapsed time, paid GPU hours and costs of the same training run on one to eight nodes
NodesEfficiencyElapsed timeGPU hoursCosts of the run
1 × NG-H100One node, no network traffic between cards100%72.0 h72.0 h€191.34$221.95
2 × NG-H100Gradient exchange via the RDMA fabric in the vRack, the private network between the nodes94%38.3 h76.6 h€203.55$236.12
4 × NG-H100Gradient exchange via the RDMA fabric in the vRack, the private network between the nodes88%20.5 h81.8 h€217.43$252.22
8 × NG-H100Gradient exchange via the RDMA fabric in the vRack, the private network between the nodes80%11.3 h90.0 h€239.18$277.45

Efficiency is the share of theoretical acceleration that remains after deducting gradient exchange. The costs are calculated as GPU hours times €2.6575$3.0827. All amounts net, plus 19% VAT.

The row with eight nodes reads like a trade-off: the run finishes after 11.3 instead of 72 hours, but costs €47.84$55.49 more than on one node. Whether this is worth it is not decided by the price, but by the question of how often you want to see a result in a day.

How long your own run takes depends on data volume and method; the catalogue does not answer this. It answers the price: you pay for the occupied GPU hours, billed to the second. A run that aborts halfway through costs half — there is neither a started day nor a minimum term.

Four identical compute nodes next to each other on a steel bench, a short connecting cable in the same arc between each pair of neighbours.
8 nodes of class NG-H100 shorten the reference run from 72 to 11.3 hours. You pay for 90.0 instead of 72 GPU hours — the gradient exchange between the nodes costs 20% efficiency.

From checkpoint to HTTPS endpoint

You provide weights or an image, and the platform turns them into an address with an access key, certificate and scaling. You are billed for the time a replica — a running copy of the model — is running, not the time the endpoint is deployed.

Scale to zero

Without requests, the endpoint shuts down its replicas and costs nothing. The first subsequent request triggers a cold start — loading the weights into the graphics memory before the first response arrives. All further requests hit a running replica. If you do not want this waiting time, set the lower limit to one — then scaling is elastic upwards and limited downwards.

  • At least zero replicas

    The first request after an idle period waits for the cold start

    Internal tools, batch processes, anything used during office hours

  • At least one replica

    No cold start time, one replica runs continuously

    Applications with users waiting for a response

  • At least two replicas

    The endpoint responds even if a fire zone fails

    Services with an availability commitment

Calculation example

An endpoint that only responds during the day

One replica on NG-L4, active for 8 hours per day over 30 days. Outside this time, the endpoint is at zero.

Class hourly rate
€0.7397$0.8581
Active hours per month
240 h
Cost with scaling to zero
€177.53$205.93
Same class continuously
€540.00$626.40
Savings compared to continuously
67%

Net amounts, Frankfurt location. The hourly rate is the same as in the instance catalogue — the platform just bills it to the second and does not charge for time without a replica.

Cold start and concurrent requests

A cold start consists of two parts: the weights are read from the local NVMe at 6 GB/s, then the runtime needs 11 seconds for CUDA context, graph construction and the first health check. Both values are in the catalogue file, so the column below can be verified and is not just a claim.

Cold start time and concurrent requests per replica for typical model sizes
Model in the endpointPrecisionWeightsClassCold startConcurrent requests
Language model, 8 billion parameters8 bit8 GBNG-L412.3 s24
Language model, 24 billion parameters8 bit24 GBNG-L40S15.0 s32
Image model, diffusionBF167 GBNG-L40S12.2 s12
Language model, 70 billion parameters8 bit70 GBNG-H10022.7 s8
Embedding model, 568 million parametersBF161.2 GBNG-L411.2 s64

Concurrent requests apply per replica with 4,096 tokens of context and an 8-bit cache; a longer context lowers the value. Requests above this limit start another replica instead of failing.

You are billed for what went over the wire

A call to a model name, with no node and no deployment. The catalogue contains only models with open weights, and they run on our own cards at locations in Germany and Finland.

The context column states the context window: how many tokens — question and previous answer combined — a model can oversee at the same time. Anything beyond that falls off the end, which is why the number is a limit and not an advertising value.

Model catalogue with task, context window, underlying instance class, measured throughput and price per million tokens
ModelTaskContextClassThroughputInput / millionOutput / million
Ministral 8B Instruct8 billion · Apache 2.0Dialogue and tool calls32,768 tokensNG-L4912 tokens/s€0.0704$0.0817€0.2816$0.3267
Mistral Small Instruct24 billion · Apache 2.0Dialogue and tool calls128,000 tokensNG-L40S1,152 tokens/s€0.1041$0.1208€0.4162$0.4828
Qwen2.5 Coder 32B32 billion · Apache 2.0Program code32,768 tokensNG-L40S648 tokens/s€0.1850$0.2146€0.7399$0.8583
Mixtral 8×7B Instruct47 billion, 13 billion active · Apache 2.0Dialogue and tool calls32,768 tokensNG-RTX60004,560 tokens/s€0.0313$0.0363€0.1251$0.1451
Llama 3.3 70B Instruct70 billion · Llama 3.3 CommunityDialogue and tool calls128,000 tokensNG-H200×84,224 tokens/s€0.2158$0.2503€0.8631$1.0012
BGE-M3568 million · MITEmbeddings8,192 tokensNG-L4179,200 tokens/s€0.0014$0.0016no output price

All amounts net, plus 19% VAT. The throughput is the aggregated value of a replica across all concurrent streams, measured with continuous batching. Embedding models return vectors — series of numbers used to compare texts — and therefore have no output price.

How the token price is calculated

We did not invent a token price, but divided the hourly rate of the card by the throughput. A replica costs an amount from the instance catalogue per hour and delivers a measurable number of tokens in that hour. The quotient is the cost rate per token; divided by the assumed utilisation, it results in the list price.

Input tokens are cheaper because in the prefill — reading the request — they run across the tensor cores and are not bound by memory bandwidth like decoding, generating the response token by token. It is set at 25% of the output price.

Why the prices vary so much between models is shown in the throughput column: a model with mixed experts — Mixture of Experts in English — activates only a fraction of its weights per token and is therefore cheaper than a dense model of the same size. It is not the parameter count that determines the price, but the number of tokens per card hour.

Calculation basis

Hourly rate
from the instance catalogue, 730-hour month
Throughput
aggregated per replica, all streams
Assumed utilisation
80%
input to output
25%
utilisation in batch mode
95%
throughput in batch mode
× 1.6

The assumed utilisation is below one hundred percent because an endpoint reserves capacity for load peaks. This reserve costs us money and is therefore included in the price, instead of being passed on to the caller as waiting time.

Batch processing

If you do not need to guarantee a response time, you submit your requests as a file and collect the result as soon as it is ready — usually within 24 hours. Without a latency commitment, we can pack more densely: utilisation increases from 80% to 95%, and larger batches raise the throughput per card by a factor of 1.6. Both together result in the discount of 47% — not a negotiated discount, but the amount we actually save.

Price per million tokens in synchronous operations and batch processing
ModelSynchronous inputSynchronous outputBatch inputBatch output
Ministral 8B Instruct€0.0704$0.0817€0.2816$0.3267€0.0371$0.0430€0.1482$0.1719
Mistral Small Instruct€0.1041$0.1208€0.4162$0.4828€0.0548$0.0636€0.2190$0.2540
Qwen2.5 Coder 32B€0.1850$0.2146€0.7399$0.8583€0.0974$0.1130€0.3894$0.4517
Mixtral 8×7B Instruct€0.0313$0.0363€0.1251$0.1451€0.0165$0.0191€0.0658$0.0763
Llama 3.3 70B Instruct€0.2158$0.2503€0.8631$1.0012€0.1136$0.1318€0.4542$0.5269
BGE-M3€0.0014$0.0016no output price€0.0007$0.0008no output price

Amounts net. A batch job is accepted as soon as it is submitted; processing begins when capacity becomes available. Results are stored as a file in the Object Storage of the project.

Rows of milled cooling fins in strict alignment, photographed vertically from above, bare copper heat pipes across them, a narrow warm streak of light on the right edge.

A cancelled run resumes from the last complete checkpoint, not at step zero: lost are the hours since this checkpoint, not the 72 of the entire run. The checkpoints are located in the Object Storage of the same location — at ENTRONYX CLOUD, neither a checkpoint nor a request leaves the EU.

Reference run 72 h · 11.3 h on 8 nodes

Where inference runs and what happens to your inputs

With an AI service, the operating location is not a footnote. Anyone who sends text to a model hands it over. The short answer: ENTRONYX CLOUD does not train on customer data, and no input leaves the EU. There is no path to a third country because no sub-processor is involved in inference. The long answer is below in sentences, not in a list of seals.

Questions about the data sovereignty of the AI platform and the corresponding answers
QuestionAnswer
Where does the inference run?Exclusively at the locations in Germany and Finland where the platform runs. A call without a location specified goes to the nearest one, never beyond. An outflow to a third country — any country outside the EU and the EEA — is therefore excluded, and outsourcing to a model provider does not take place.
What happens to the inputs?They are kept in memory for the duration of the request and then discarded. No writing to storage, no storing in a cache, no training on your data — neither for our models nor for those of third parties. This commitment is also stated in the data processing agreement.
What is logged?Time, model name, token count, response time and status code — the details from which the invoice is generated. Not the content of the request and not the content of the response.
Is there any transfer to third parties?No. There is no sub-processor for inference. The models are stored as open weights on our own storage media and run on our own hardware.
Who is the legal operator of the platform?A German company based in Germany, with no parent company outside the EU. Requests for disclosure under foreign law are therefore void; the mutual legal assistance procedure is authoritative.
What is in the data processing agreement?The agreement under Article 28 GDPR names the locations individually, excludes processing outside the EU and lists the technical and organisational measures. It can be viewed before the contract is concluded.
Are the models verifiable?The weights are open and their licence is noted in the catalogue. The model version and checksum of each endpoint are in the console; a model change gets a new name and never silently replaces the old one.

These details are part of the data processing agreement. If the contract text deviates from this table, the contract applies — please report the discrepancy to us, as it is then an error on this page.

Inference locations

4 locations in Germany and Finland. The complete model catalogue is served at 3 of them; the others run the models whose instance class is listed there. A request without a location specified will always remain within this list. PUE is the energy efficiency metric of a data centre — the total power consumption divided by the power of the computing equipment; 1.0 would be lossless.

  • Frankfurt am Mainfra1

    Location Rhein-Main I · Germany

    PUE 1.14 · 3 ms

    • ISO 27001
    • ISO 50001
    • BSI C5:2020
    • EN 50600 VK4
    • TÜV Rheinland

    Complete model catalogue

  • Frankfurt am Mainfra2

    Location Rhein-Main II · Germany

    PUE 1.09 · 3 ms

    • ISO 27001
    • ISO 50001
    • BSI C5:2020
    • EN 50600 VK4

    Complete model catalogue

  • Munichmuc1

    Location Isar · Germany

    PUE 1.11 · 5 ms

    • ISO 27001
    • ISO 50001
    • BSI C5:2020
    • EN 50600 VK4
    • Suitable for critical infrastructure
  • Helsinkihel1

    Location Uusimaa · Finland

    PUE 1.08 · 21 ms

    • ISO 27001
    • ISO 50001
    • BSI C5:2020
    • EN 50600 VK4

    Complete model catalogue

What we do not offer

No passing on to third-party model APIs
We do not operate a proxy system that forwards requests to an American provider and passes them off as our own offering. What is in the catalogue runs on our cards.
No closed weights
We cannot run a model on our own hardware if we are not allowed to deliver its weights. Therefore, the catalogue only contains models with an open licence.
No silent model changes
A model name remains tied to a weight version. New versions appear as a separate name; old ones are phased out with an announced notice period.

Configure a service or rent the card directly

The configurator guides you through the working environment, training job and endpoint, showing the same hourly rate as in the tables above. If you need the hardware without the platform on top, go directly to the GPU catalogue; the complete overview of all items is in the price list.

Accelerators from
€0.7397$0.8581
Workspace from
€0.0567$0.0658
Output from
€0.1251$0.1451
Inference locations
4