ComputeOperational
Bare Metal, cloud instances, GPU nodes, Kubernetes workers
- 90 days
- 99,962 %
- Commitment
- 99,95 %
All hypervisor clusters within target range. No provisioning backlog.
This page is fed by the same measurement that determines SLA service credits at ENTRONYX CLOUD — the SLA is our availability commitment. We also publish values here that cost us money. An incident is a disruption that we track publicly. It appears as soon as three independent checkpoints report a deviation.
Every status is indicated by a symbol and text, not just by colour. One bar in the availability band corresponds to one day; its height decreases with the severity.
2 of 6 services deviate from normal operations, 1 of 5 locations are affected. Details per service below.
4 operational · 1 affected
Normal operations
Block Storage: increased write latency on pool 3 and 7
Normal operations
Normal operations · cooling circuit 2 in manual mode since 06:12, without impact
Normal operations
Measured from 14 external checkpoints outside our network. A day is considered disrupted as soon as at least three checkpoints report a deviation simultaneously.
Bare Metal, cloud instances, GPU nodes, Kubernetes workers
All hypervisor clusters within target range. No provisioning backlog.
Object Storage, Block Storage, backup targets, cold archive
fra2: increased write latency on two Block Storage pools after replacing a controller. Read path unaffected.
Backbone, transit, peering, vRack, DDoS filter
18.4 Tbit/s capacity, peak load at 41%. All rings closed.
api.entronyx.cloud, S3 endpoints, Terraform provider
P95 response time at 118 ms across all endpoints.
Web interface, billing, ticket system, two-factor authentication
Maintenance window: migration of the invoice view to the new document archive API. Login and ticket system are running.
Authoritative zones, resolvers, reverse DNS
12 Anycast locations reachable. No location removed from the cluster.
One bar per day. The height of the bar decreases with the severity — the meaning does not rely solely on the colour.
A maintenance window is a previously announced period during which we work on a system. We announce work without expected interruption five calendar days in advance, and work with expected interruption 14 days in advance. Upon request, we can postpone individual windows on your dedicated resources — systems used only by you — by up to 14 days.

2 September 2026, 02:00–04:00 MESZ
Customer panel · all regions
The document archive is being migrated to a separate service with full-text search across all invoices and service credits. The migration includes 4.1 million documents.
Expected impact: Invoice view and PDF download will be unavailable for up to 40 minutes. API, Compute, Network and Storage are not affected.
9 September 2026, 01:00–05:00 MESZ
hel1 · fire section A
Mandatory test according to EN 50600 with full load transfer to the emergency power system. The UPS bridges the switching process.
Expected impact: No expected impact on running systems. Servers with only one connected power supply unit are theoretically at risk during the switchover — affected customers have been informed individually.
16 September 2026, 23:00 – 17 September 2026, 03:00 MESZ
Backbone · North route (hel1 ↔ ber1 ↔ fra1)
Installation of additional transponders and switching of wavelengths. The ring will be opened on one side for this purpose, traffic will run via the opposite direction.
Expected impact: Latency between hel1 and fra1 will increase by up to 9 ms for the duration of the window. No packet loss expected, no interruption.

Each entry first names the affected zone — the area to which the disruption was confined, i.e. a location, a service or a transit route, meaning a connection to an external network. This is followed by the timeline, impact, cause, resolution and whether a service credit was triggered. We write the log exactly as we keep it internally — including the parts where we ourselves were the reason for the delay.

fra2 · Block Storage pool 3, 7 and 11 · around 2,900 volumes
Write operations on the affected pools reached latencies up to 4,200 ms, at peak database transactions timed out. Read accesses remained consistently below 3 ms. No data loss: the journal layer retained all confirmed write operations, the checksum runs after the resolution were error-free.
A firmware update on 96 NVMe shelves activated a background garbage collection whose default parameters are set too aggressively for our write profiles. The manufacturer did not mark the change as behaviour-relevant in the release notes; our preliminary testing in the lab ran with a load profile that did not trigger the condition.
Rollback of the firmware on all affected shelves in waves of 8 shelves each, followed by a controlled rebuild of the pools. From 19:40, latencies were back below 12 ms, by 21:52 all rebuilds were completed.
Honest retrospective: two things went wrong here that have nothing to do with the manufacturer. Firstly, we did not have a canary ring for storage firmware — the rollout went to the entire pool group in one step. Since 18 August, firmware changes run via a canary of 4% of the shelves with 72 hours of observation. Secondly, we only listed this incident on the status page at 14:48, 41 minutes after the first internal alert; customers wrote tickets during this time and received no response. The threshold for a status entry is now at 10 minutes of confirmed deviation and is triggered automatically from the alert, no longer decided manually.
Service credit of 25% of the monthly fee for all affected volumes was issued without request on the 09/2026 invoice.
Alert: P99 write latency pool 3 exceeds 500 ms.
Status entry published, cause still unclear.
Connection with the firmware rollout confirmed, rollout stopped.
Start of the rollback, first wave of 8 shelves.
Latencies back in the normal range, rebuild running.
All pools synchronous, incident closed.
hel1 · outgoing transit via one of four upstreams
For destinations reached via this upstream, packet loss was between 4% and 11%. Destinations in the Baltics and Northern Europe were predominantly affected. Traffic within the ENTRONYX backbone as well as via DE-CIX, BCIX and FICIX was not affected at any time.
The upstream started announced maintenance on a linecard chassis earlier than communicated. The BGP session flapped instead of tearing down cleanly, causing our route selection to oscillate between two paths for twelve minutes.
Manual reduction of the Local Preference for the affected upstream, rerouting to the remaining three transit contracts. Traffic subsequently ran without loss; the session was brought back in a controlled manner on 3 July at 11:00.
Route selection now automatically dampens sessions if more than two state changes occur within five minutes (BGP damping with adjusted thresholds).
Below the service credit threshold: network availability in July was 99.994%, which is above the committed value.
Synthetic probes report loss on paths via AS path 3.
Status entry published.
Local Preference lowered, traffic is being rerouted.
Probes green again, incident closed.
api.entronyx.cloud · all regions · POST, PUT, PATCH, DELETE
All write API calls were answered with HTTP 502. Read calls and the S3 data path remained available. The customer panel subsequently showed empty resource lists because it accesses a write session endpoint when loading. Existing instances, networks and storage continued to run unchanged — only the control plane was affected.
A configuration change to the rate limiter was rolled out with a threshold of 0 instead of 3,000 requests per minute. The change passed the check because our schema interpreted the value 0 as “unlimited”, but the limiter itself interpreted it as “nothing allowed”.
Rollback of the configuration via the previous revision state. Recovery took 90 seconds from the detection of the cause.
The configuration schema no longer recognises 0; unlimited is explicitly listed as null. In addition, a progressive rollout across three rings with automatic cancellation at an error rate above 1% has applied to the control plane since 25 June.
14 minutes corresponds to 99.968% API availability in June. The value is above the commitment of 99.95%, so a service credit was not triggered.
Error rate of write endpoints jumps to 100%.
Status entry published, control plane affected.
Cause identified, rollback started.
Error rate back to background noise, incident closed.
Authoritative DNS · 11 of 12 Anycast locations
Changes to DNS records were applied at eleven locations with a delay of up to 40 minutes. Existing records were answered correctly throughout, there were no NXDOMAIN responses and no signature errors.
During the scheduled rollover of the Zone Signing Key, the re-signing of 61,000 zones ran on the same nodes as the transfer service. CPU saturation slowed down the transfers.
Signing moved to dedicated nodes and the rollover continued in batches of 5,000 zones.
No service credit: name resolution itself was not disrupted at any time.
Transfer delay detected at eleven locations.
Status entry published.
Signing moved to separate nodes.
All locations synchronous, incident closed.
We keep older incidents for 36 months and make them available on request. Security-related incidents do not appear in this public log; affected customers are informed directly according to the procedure in Security and compliance.
You can be notified about new incidents and planned maintenance — either for all services or only for the locations and products you actually use. You make this selection in the customer panel under “Notifications”.
For “critical” level incidents, we also send a message to all registered technical contacts, regardless of the selected settings.
Atom feed — a machine-readable news stream — covering all incidents and maintenance windows, filterable by region and service.
/status.atom
A webhook is an automatic request to your address (HTTP POST) on every status change. The payload is signed with HMAC-SHA256 so you can verify that the message comes from us.
Panel → Notifications
Planned maintenance windows as an iCalendar subscription — a calendar that your team calendar system integrates and continuously fetches.
/status.ics
Before you write a ticket: this page clarifies whether we already know about the disruption. If everything is running within the target range on our end but not on yours, a resource ID, timestamp with time zone, and an excerpt from the log will help most.