Skip to content

GPU & KI

6 card types in 10 configurations: from the inference card that lets a fully trained model respond, to the 8 GPU node with 1,128 GB graphics memory. Billed by the second at the hourly rate and capped at the monthly price: a three-day training run costs three days.

Drivers, CUDA toolchain (NVIDIA's programming platform for the GPU) and container runtime are pre-installed. The local NVMe, the fast flash storage of the instance, is in the same chassis as the card — datasets are not read over the network.

Entry
€214.00$248.24
NG-RTX4000, net per month · €0.2932$0.3401 per hour
Configurations
10
6 card types, RTX PRO 4000 to H200
VRAM in the node
1.128 GB
8 × H200 via NVLink
Locations
4
with GPU capacity

Three questions, and the card is decided

Does the model fit into memory, do multiple cards need to talk to each other, and what does an hour cost? The decision needs nothing more. Everything else on this page is evidence.

  1. Question 1

    How much memory does the card have?

    24 – 141 GB

    VRAM — the memory on the card itself — is the first limit. What doesn't fit, won't run. In the largest node, 8 cards add up to 1,128 GB.

  2. Question 2

    How fast do the cards talk to each other?

    900 GB/s

    This is what NVLink — the direct connection between two GPUs in the same chassis — carries per card in the NG-H200×8 node. All other lines communicate via PCIe, the motherboard's slot bus, and are thus about an order of magnitude lower.

  3. Question 3

    What does an hour cost?

    from €0.2932$0.3401

    The range ends at €42.00$48.72 per hour for the largest node. Broken down to a single card, it starts at €0.2932$0.3401 per GPU hour; a run over 72 hours thus starts at €21.11$24.49 net.

Two other words appear again and again on this page. Throughput means how many calculation steps run per second — specified in TFLOPS, according to the manufacturer here between 24 and 989 per card, measured on a different basis. The TDP is the waste heat a card emits in continuous operation; it decides at which location a node may stand. The L4 stays at 72 watts and manages without an additional power connection.

ENTRONYX CLOUD carries each card in several expansion stages. The choice is therefore made twice: first the card type, then the number in the node.

Where this explicitly falls short

  • The smallest card does not train. With 24 GB and 24 TFLOPS, it answers well and learns poorly. Fine-tuning — retraining a finished model on your own data — waits there at the memory bandwidth.
  • Four cards are not a cluster. Without NVLink, coordination between the cards runs via PCIe. For communication-heavy training across multiple GPUs, this is the bottleneck, not the computing power.
  • A model that does not fit will not get faster. If you swap out, you lose more time on the bus than the larger card brings in computing power. First the memory, then the throughput.

Which card fits your model

The GPU never stands alone. Every row brings dedicated processor cores, memory and local NVMe — otherwise the data path becomes a bottleneck before the card is fully utilised. 6 card types thus result in 10 configurations.

GPU configurations with graphics card, computing power, processor, memory, local NVMe, traffic, hourly rate and monthly cap

  • NG-RTX4000Entry1 × RTX PRO 4000 · 24 GB GDDR7

    €0.2932$0.3401per hour

    Intel Core i5-13500 (14 cores, 20 threads)

    TFLOPS
    24
    vCPU
    20
    Memory
    64 GB DDR4
    Local NVMe
    2 × 512 GB NVMe
    Traffic
    unlimited
    Cap per month
    €214.00$248.24
  • NG-L4Inference1 × L4 · 24 GB GDDR6

    €0.7397$0.8581per hour

    AMD EPYC (dedicated)

    TFLOPS
    121
    vCPU
    8
    Memory
    48 GB
    Local NVMe
    400 GB
    Traffic
    40 TB
    Cap per month
    €540.00$626.40
  • NG-L40S1 × L40S · 48 GB GDDR6

    €1.3808$1.6017per hour

    AMD EPYC (dedicated)

    TFLOPS
    362
    vCPU
    16
    Memory
    96 GB
    Local NVMe
    800 GB
    Traffic
    50 TB
    Cap per month
    €1,008.00$1,169.28
  • NG-RTX60001 × RTX PRO 6000 · 96 GB GDDR7

    €1.6425$1.9053per hour

    Intel Xeon Gold 5412U (24 cores, 48 threads)

    TFLOPS
    110
    vCPU
    48
    Memory
    256 GB DDR5 ECC
    Local NVMe
    2 × 960 GB NVMe
    Traffic
    unlimited
    Cap per month
    €1,199.00$1,390.84
  • NG-RTX6000-5121 × RTX PRO 6000 · 96 GB GDDR7

    €2.4644$2.8587per hour

    Intel Xeon Gold 5412U (24 cores, 48 threads)

    TFLOPS
    110
    vCPU
    48
    Memory
    512 GB DDR5 ECC
    Local NVMe
    2 × 1.92 TB NVMe
    Traffic
    unlimited
    Cap per month
    €1,799.00$2,086.84
  • NG-RTX6000-7681 × RTX PRO 6000 · 96 GB GDDR7

    €3.1493$3.6532per hour

    Intel Xeon Gold 5412U (24 cores, 48 threads)

    TFLOPS
    110
    vCPU
    48
    Memory
    768 GB DDR5 ECC
    Local NVMe
    4 × 3.84 TB NVMe
    Traffic
    unlimited
    Cap per month
    €2,299.00$2,666.84
  • NG-H100Training1 × H100 · 80 GB HBM

    €2.6575$3.0827per hour

    AMD EPYC (dedicated)

    TFLOPS
    989
    vCPU
    32
    Memory
    256 GB
    Local NVMe
    3.2 TB
    Traffic
    80 TB
    Cap per month
    €1,940.00$2,250.40
  • NG-H100×22 × H100 · 160 GB HBM

    €5.3151$6.1655per hour

    2× AMD EPYC (dedicated)

    TFLOPS
    1,978
    vCPU
    64
    Memory
    512 GB
    Local NVMe
    6.4 TB
    Traffic
    100 TB
    Cap per month
    €3,880.00$4,500.80
  • NG-H100×44 × H100 · 320 GB HBM

    €10.6438$12.3468per hour

    2× AMD EPYC (dedicated)

    TFLOPS
    3,956
    vCPU
    128
    Memory
    1 TB
    Local NVMe
    12.8 TB
    Traffic
    140 TB
    Cap per month
    €7,770.00$9,013.20
  • NG-H200×8Cluster8 × H200 · 1,128 GB HBM

    €42.0000$48.7200per hour

    2× AMD EPYC (dedicated)

    TFLOPS
    7,912
    vCPU
    128
    Memory
    2 TB
    Local NVMe
    30.72 TB
    Traffic
    200 TB
    Cap per month
    €30,660.00$35,565.60

TFLOPS are manufacturer specifications with different measurement bases and therefore not directly comparable between cards. The graphics memory is the primary factor for selection.

Net amounts, location Frankfurt am Main. The hourly rate is the monthly price divided by 730 hours. Billing is per second at the hourly rate, capped at the monthly price — you will never be charged for more than 730 hours a month.

Four accelerator cards side by side in an open compute node, contact strips on the slots, bundled power cables on the right.
An accelerator card always occupies its own slot with full bandwidth. There are no shared cards here — what is in the row is available to the instance alone.
Bundle of copper cables of equal length between two server racks, each end labelled, behind them the backs of two compute nodes.

Within a node, NVLink connects the 8 cards at 900 GB/s per GPU. Between two nodes, only the network carries the load. ENTRONYX CLOUD sets up a dedicated RDMA fabric in the vRack for this — a private network where a card writes directly to the memory of the other side. Gradients therefore do not run over standard TCP.

NG-H200×8 · NVLink in the chassis, RDMA in between

Which card for which task

The most expensive card is rarely the right one. What matters is whether your model fits into memory, whether you are training or responding, and whether multiple GPUs need to talk to each other.

NG-L4

Inference and transcoding

The L4 is a single-slot card with 72 watts and a seventh-generation encoder. Its home is constantly running services: responses from quantised language models up to around 13 billion parameters, image generation, embedding calculation, and video transcoding including AV1.

VRAM
24 GB
TFLOPS (manufacturer specification)
121
Price / hour
€0.7397$0.8581

Not for training. If you fine-tune here, you will be waiting on memory bandwidth.

NG-L40S

Fine-tuning and rendering

48 GB of graphics memory is enough for LoRA and QLoRA adaptations on models up to around 30 billion parameters, without the detour of offloading to main memory. The RT cores make the same card the best choice for render farms and simulation previews.

VRAM
48 GB
TFLOPS (manufacturer specification)
362
Price per 100 TFLOPS/h
€0.38$0.44

No NVLink bridge: multiple L40S communicate via PCIe. This is the bottleneck for communication-heavy multi-GPU training.

NG-RTX6000

Inference and fine-tuning of medium-sized models

96 GB GDDR7 with ECC — the largest graphics memory among the single cards in the catalogue. The RTX PRO 6000 Blackwell holds 8-bit quantised models up to the mid-double-digit billion range including KV cache on a single card and handles LoRA and QLoRA adaptations without offloading. The H100 class with HBM remains responsible for bandwidth-bound training.

VRAM
96 GB
Local NVMe
1.92 TB
72-hour run
€118.26$137.18
NG-H100

Large-scale training

Hopper with 80 GB HBM3, 3.35 TB/s bandwidth, and Transformer Engine. The FP8 path doubles the throughput compared to BF16, provided the training script uses it. For models that fit into one GPU but would take weeks instead of days, this is the most economical single card in the catalogue.

VRAM
80 GB
TFLOPS (manufacturer specification)
989
72-hour run
€191.34$221.95
NG-H200×8

Distributed cluster training

Eight H200s in one node, connected via NVLink with 900 GB/s per GPU — not via Ethernet. A total of 1,128 GB graphics memory, enough for models that could not even be loaded on a single card. Tensor and pipeline parallelism run within the chassis, without a network hop.

Total VRAM
1.128 GB
Main memory
2 TB
Local NVMe
30.72 TB

Only bookable after a capacity check. We link multiple nodes via a dedicated RDMA fabric in the vRack.

Calculate training costs

Configuration, location, hours and nodes: the calculator uses the same rates and the same monthly cap as the configurator. Below is a calculated run over 72 hours for each configuration — net, in Frankfurt am Main, with instance, local NVMe and traffic.

Configuration

1 × RTX PRO 4000 · 24 GB graphics memory per node

Location

Reference location — the catalogue price applies without a factor.

Runtime per month

1 h730 h

72 hours · 3 days

Nodes

1 Card · 24 GB graphics memory · 1 to 99 nodes

A training run over 72 hours, calculated

Three days is a realistic window for fine-tuning with subsequent evaluation.

Cost of a 72-hour run per configuration, plus the price per individual GPU hour and per 100 TFLOPS
ConfigurationGPUsTotal VRAMPrice / hourper GPU hourper 100 TFLOPS/h (manufacturer specification, see note below the table)72 hours
NG-RTX4000NVIDIA RTX PRO 4000 Blackwell SFF124 GB€0.2932$0.3401€0.2932$0.3401€1.22$1.42€21.11$24.49
NG-L4NVIDIA L4124 GB€0.7397$0.8581€0.7397$0.8581€0.61$0.71€53.26$61.78
NG-L40SNVIDIA L40S148 GB€1.3808$1.6017€1.3808$1.6017€0.38$0.44€99.42$115.33
NG-RTX6000NVIDIA RTX PRO 6000 Blackwell Max-Q196 GB€1.6425$1.9053€1.6425$1.9053€1.49$1.73€118.26$137.18
NG-RTX6000-512NVIDIA RTX PRO 6000 Blackwell Max-Q196 GB€2.4644$2.8587€2.4644$2.8587€2.24$2.60€177.44$205.83
NG-RTX6000-768NVIDIA RTX PRO 6000 Blackwell Max-Q196 GB€3.1493$3.6532€3.1493$3.6532€2.86$3.32€226.75$263.03
NG-H100NVIDIA H100 SXM180 GB€2.6575$3.0827€2.6575$3.0827€0.27$0.31€191.34$221.95
NG-H100×2NVIDIA H100 SXM2160 GB€5.3151$6.1655€2.6576$3.0828€0.27$0.31€382.69$443.92
NG-H100×4NVIDIA H100 SXM4320 GB€10.6438$12.3468€2.6610$3.0868€0.27$0.31€766.35$888.97
NG-H200×8NVIDIA H200 SXM81,128 GB€42.0000$48.7200€5.2500$6.0900€0.53$0.61€3,024.00$3,507.84
TFLOPS are manufacturer specifications with different measurement bases and therefore not directly comparable between cards — the column per 100 TFLOPS provides context, it does not decide.

Two columns deserve a second look. The first is per GPU hour. The eight-GPU node costs €5.2500$6.0900 per card there and is thus only slightly above a single H100 with €2.6575$3.0827. It provides 76% more graphics memory per card and NVLink between all eight.

The second is per 100 TFLOPS. Measured against the manufacturer's specification for computing power, the H100 is the cheapest row in the table at €0.27$0.31. The L40S is the best compromise without HBM at €0.38$0.44 — the stacked high-performance memory that only the H100 class carries. The RTX PRO 6000, on the other hand, buys graphics memory, not throughput at €1.49$1.73: 96 GB for models that would otherwise require two cards. If you need computing power instead of memory, you should know this order before reaching for the largest number.

The path remains the same in both cases. First check whether the model fits into the memory. Then compare the price per computing power. And only then select the card by name.

Calculation basis

Term in the example
72 h
Billing cycle
per second
Location
Frankfurt am Main (fra1)
Price factor
1.00
Cap per month
730 h
Included
Instance, local NVMe, traffic

Billing is per second at the hourly rate, capped at the monthly price — you will never be charged for more than 730 hours a month. The hourly rate is calculated on 730 hours, the average month. A month with 31 days has 744 — we do not charge for the additional hours.

Select pricing unit

Monthly price is also the cap for hourly billing

GPU configurations with graphics memory, computing power and price optionally per hour or per month
ConfigurationGPUsTotal VRAMTotal TFLOPS72 hPrice / month
NG-RTX4000NVIDIA RTX PRO 4000 Blackwell SFF124 GB24€21.11$24.49€214.00$248.24
NG-L4NVIDIA L4124 GB121€53.26$61.78€540.00$626.40
NG-L40SNVIDIA L40S148 GB362€99.42$115.33€1,008.00$1,169.28
NG-RTX6000NVIDIA RTX PRO 6000 Blackwell Max-Q196 GB110€118.26$137.18€1,199.00$1,390.84
NG-RTX6000-512NVIDIA RTX PRO 6000 Blackwell Max-Q196 GB110€177.44$205.83€1,799.00$2,086.84
NG-RTX6000-768NVIDIA RTX PRO 6000 Blackwell Max-Q196 GB110€226.75$263.03€2,299.00$2,666.84
NG-H100NVIDIA H100 SXM180 GB989€191.34$221.95€1,940.00$2,250.40
NG-H100×2NVIDIA H100 SXM2160 GB1,978€382.69$443.92€3,880.00$4,500.80
NG-H100×4NVIDIA H100 SXM4320 GB3,956€766.35$888.97€7,770.00$9,013.20
NG-H200×8NVIDIA H200 SXM81.128 GB7,912€3,024.00$3,507.84€30,660.00$35,565.60

The column for 72 hours remains unchanged — it always shows the amount actually incurred for a three-day run.

Hourly, monthly or with a fixed term

Three ways to the same card. Which one is the cheapest depends solely on utilisation — not on the model and not on the location.

Monthly price per configuration by term, plus the savings of one year with a 24-month commitment
Configurationmonthly12 months −8%24 months −14%Savings / year
NG-RTX4000€214.00$248.24€196.88$228.38€184.04$213.49€359.52$417.00
NG-L4€540.00$626.40€496.80$576.29€464.40$538.70€907.20$1,052.40
NG-L40S€1,008.00$1,169.28€927.36$1,075.74€866.88$1,005.58€1,693.44$1,964.40
NG-RTX6000€1,199.00$1,390.84€1,103.08$1,279.57€1,031.14$1,196.12€2,014.32$2,336.64
NG-RTX6000-512€1,799.00$2,086.84€1,655.08$1,919.89€1,547.14$1,794.68€3,022.32$3,505.92
NG-RTX6000-768€2,299.00$2,666.84€2,115.08$2,453.49€1,977.14$2,293.48€3,862.32$4,480.32
NG-H100€1,940.00$2,250.40€1,784.80$2,070.37€1,668.40$1,935.34€3,259.20$3,780.72
NG-H100×2€3,880.00$4,500.80€3,569.60$4,140.74€3,336.80$3,870.69€6,518.40$7,561.32
NG-H100×4€7,770.00$9,013.20€7,148.40$8,292.14€6,682.20$7,751.35€13,053.60$15,142.20
NG-H200×8€30,660.00$35,565.60€28,207.20$32,720.35€26,367.60$30,586.42€51,508.80$59,750.16

What starts where straight away — and what we check first

GPUs are the only product in this catalogue for which we do not guarantee immediate availability. We therefore state the situation per configuration and location in advance. Where it says “on request”, the link leads directly to the request — configuration and location are already filled in.

immediately
always available, start without review
on request
available based on capacity, the status per card is listed below
not offered
not available at this location

Availability of each GPU configuration per location, as of 17 September 2026

Situation per card type, as of 17 September 2026 — one row per card, valid for all expansion stages of this card: immediate availability, reservation lead time and quota per project
CardOn-demandReservationQuota
NVIDIA RTX PRO 4000 Blackwell SFFBase row NG-RTX4000continuously availablenot required8 GPUs
NVIDIA L4Base row NG-L4continuously availablenot required8 GPUs
NVIDIA L40SBase row NG-L40Scontinuously availablepossible from 30 days4 GPUs
NVIDIA RTX PRO 6000 Blackwell Max-QBase row NG-RTX6000usually availablefrom 30 days, 2 working days lead time2 GPUs
NVIDIA H100 SXMBase row NG-H100subject to capacityfrom 30 days, 5 working days lead time2 GPUs
NVIDIA H200 SXMBase row NG-H200×8only after reviewfrom 30 days, 10 working days lead timeon request
Query and reserve capacity
# Query free capacity per model and location
entronyx gpu availability --model h100 --region fra1
# MODEL    LOCATION  FREE  RESERVABLE    NEXT SLOT
# h100     fra1         3            12  now
# h200x8   fra1         0             2  on request
 
# Bindingly reserve capacity for 30 days
entronyx gpu reservation create \
  --flavor ng-h100 --region fra1 --count 4 --term 30d
 
# Notification as soon as an H200 node becomes available
entronyx gpu watch --flavor ng-h200x8 --region fra2 --notify webhook
The availability query is not rate-limited and may run in deployment scripts. Reservations appear as a separate item on the invoice.
Two hands insert an accelerator card into the slot of an open server chassis, next to it the mounting frame and the screwdriver.
For every installed card, there are 8 to 48 dedicated EPYC cores. The ratio increases with the card so that preprocessing does not slow down the GPU — no node is equipped more sparsely at ENTRONYX CLOUD.

Which card is where

GPU nodes require more power per rack unit and different cooling than a standard chassis. Therefore, not every configuration is available at every location — and the largest one only where liquid cooling is installed.

  • Frankfurt am Mainfra1

    Location Rhein-Main I

    RTX PRO 4000 Blackwell SFF · L4 · L40S · RTX PRO 6000 Blackwell Max-Q · H100 SXM · H200 SXM

    10 of 10 configurations · PUE 1.14 · 3 ms

    Complete range

  • Frankfurt am Mainfra2

    Location Rhein-Main II

    RTX PRO 4000 Blackwell SFF · L4 · L40S · RTX PRO 6000 Blackwell Max-Q · H100 SXM · H200 SXM

    10 of 10 configurations · PUE 1.09 · 3 ms

    Complete range

  • Münchenmuc1

    Location Isar

    RTX PRO 4000 Blackwell SFF · L4 · L40S

    3 of 10 configurations · PUE 1.11 · 5 ms

  • Helsinkihel1

    Location Uusimaa

    RTX PRO 4000 Blackwell SFF · L4 · L40S · RTX PRO 6000 Blackwell Max-Q · H100 SXM · H200 SXM

    10 of 10 configurations · PUE 1.08 · 21 ms

    Complete range

3 of the 4 locations carry all 10 configurations: Location Rhein-Main I, Location Rhein-Main II, Location Uusimaa. The rest carry the cards that their power density allows. PUE is the ratio of total power to computing power — the closer to 1.0, the less goes into cooling and losses. Which configuration starts immediately at which location can be found in the availability matrix.

Training data without third-country transfer

All 4 GPU locations are within the scope of the GDPR, spread across Germany and Finland. If you compute here, your training data is processed without third-country transfers. The data processing agreement according to Art. 28 GDPR is part of the main contract and not an additional form.

Data processing agreement (DPA) Contract text and sub-processors

4Locations, all without third-country transfer

Prepared, but not prescriptive

The standard image includes drivers, the CUDA toolchain and a container runtime. Everything on top is yours — we do not dictate any framework or version.

Driver
NVIDIA 570 LTS
Data centre branch, maintained monthly
CUDA
12.8 and 12.4
switchable via modules
Container
NVIDIA Container Toolkit
Docker and Podman, --gpus all
Operating system
Ubuntu 24.04 LTS
also Debian 13 and Rocky 10
First steps on a GPU instance
# Drivers, CUDA and container runtime are pre-installed in the image
nvidia-smi --query-gpu=name,memory.total,driver_version --format=csv
 
# Single GPU: container with access to the local NVMe
docker run --rm --gpus all \
  -v /mnt/nvme/datasets:/data \
  -v /mnt/nvme/checkpoints:/ckpt \
  nvcr.io/nvidia/pytorch:25.06-py3 \
  python train.py --data /data --out /ckpt --precision bf16
 
# Eight GPUs in one node: NCCL runs over NVLink, not over the network
torchrun --standalone --nproc_per_node=8 train.py \
  --precision fp8 --grad-accum 4 --checkpoint-every 500
 
# Observe usage during the run
nvidia-smi dmon -s pucm -d 5
The local NVMe is mounted at /mnt/nvme. It survives reboots, but not the deletion of the instance — checkpoints also belong in the Object Storage.

Not just language models

The L4 decodes and encodes up to eight 4K streams simultaneously in AV1, the royalty-free video format. For media libraries and live transcoding, this makes it the cheapest line in the catalogue. The L40S also brings RT cores — compute units for ray tracing — for Blender, Omniverse and Unreal render farms. Both cases pay off per hour very differently than training.

from · per month, net€214.00$248.24

Configure GPU instance

Configure GPU instance or reserve capacity

RTX PRO 4000, L4 and L40S are continuously available and start without review. You purchase RTX PRO 6000, H100 and H200 based on capacity at the selected location; a reservation of 30 days or more secures the card.

from
€214.00$248.24
VRAM in the node
1.128 GB
72-hour run from
€21.11$24.49
Locations
4