Skip to main content

Stable Release Version v3.1.9

· 5 min read

Version v3.1.9 brings customer support inside the panel. The headline is a complete Support Ticket System: customers open tickets against a department, attach the exact resource they are having trouble with, and follow the conversation in real time, while your team works the queue against per-priority SLA targets with internal notes the customer never sees. Around it, Windows guests gain first-class Hyper-V enlightenment settings and Proxmox Windows support, imported virtual machines can override their cloud-init, every outbound email now carries a proper sender name, and the admin storage lists paginate the way they always should have.

Support Ticket System

A full multi-tenant helpdesk now ships with the platform - no external tool, no separate login for your customers.

  • [Feature] Tickets, Departments and Priorities - Customers open tickets from the user panel against a department you define, at one of four priorities (low, normal, high, urgent). Each ticket carries a sequential ticket number and moves through a clear lifecycle - open, pending, resolved, closed - with resolved tickets auto-closing after a configurable window and a reopen grace period during which a customer reply brings the ticket back to life. Departments can be assigned their own set of admins so the right team sees the right queue.
  • [Feature] SLA Policies with Pause-Aware Timers - Define first-response and resolution targets per department and priority. The clock is honest: a ticket moved to pending (waiting on the customer) pauses its resolution timer and resumes it on the next reply, so time spent waiting on the customer never counts against your team. Admins get an SLA reports view with attainment computed from the actual due columns, not cron-timing-dependent flags, so the numbers hold up.
  • [Feature] Threaded Replies and Staff-Only Internal Notes - Both sides reply on one thread. Staff can additionally post internal notes that are never rendered to the customer - the visibility rule is a fail-closed allowlist, so a note can never leak into a customer-facing view or payload. Internal notes are even allowed on closed tickets, for post-mortems.
  • [Feature] Attach the Resource the Ticket Is About - A ticket can link the exact instance, volume, Kubernetes cluster, managed database, load balancer, object-storage bucket or VPC it concerns. The picker only ever shows resources the account actually owns, so your team lands on the right server instead of asking "which VM?".
  • [Feature] Virus-Scanned Attachments with In-Browser Preview - Customers and staff can attach files up to a configurable size limit. Every upload is scanned with ClamAV before it is accepted (with an admin-controlled fail-open/fail-closed policy), and images and PDFs render inline in the browser instead of forcing a download. Attachments are stored on local disk or an S3-compatible bucket of your choosing, configured from the support settings page.
  • [Feature] Notifications and Real-Time Updates - Ticket events fan out to your team over Slack, Discord, Telegram or a generic webhook, and the ticket thread updates live over WebSocket - new replies, status changes and SLA breaches appear without a refresh. Delivery is best-effort per channel, so one bad webhook can never block an email or another channel.
  • [Feature] Support Reaches Everyone, Even Suspended Accounts - The customer support surface stays reachable for suspended and unbilled accounts by design - support is how a customer disputes a bill - while an admin kill switch can disable the entire customer-facing support system in one setting when you need to. Team members respect granular permissions: separate view and manage capabilities gate the queue for subusers.

Windows Guests

  • [Feature] Hyper-V Enlightenment Settings - KVM Windows guests can now be given the Hyper-V enlightenments Windows expects for stable timers and performance. Settings apply as hypervisor-wide defaults with a per-instance override, exposed on both the user and admin panels, and are shown on every KVM instance running Windows whether it was installed from an ISO or imported.
  • [Feature] Proxmox Windows Support - On Proxmox hypervisors a single Windows-guest switch drives the VM's OS type, so Proxmox derives the correct enlightenment set itself rather than requiring the KVM-style detail.

Instances

  • [Feature] Cloud-Init Override for Imported VMs - Instances now carry an optional per-instance cloud-init override, so an imported virtual machine can be handed the exact cloud-init configuration it needs instead of inheriting only the platform default.

Email and Admin

  • [Fix] Outbound Mail Sender Name - Every email the platform sent used a bare from-address with no display name, so mail arrived from a nameless sender (most visibly on the new support notifications). All outbound mail now carries the sender name from your mail settings - falling back to the application name - across all 43 mailables, from support and authentication through billing, backups, Kubernetes and Let's Encrypt.
  • [Fix] Admin Storage Pagination - The admin Storages and Backup Storages lists advanced the page in the URL but kept showing page one. Both pages now refresh their rows when the page changes, matching every other admin list.

Stable Release Version v3.1.7

· 13 min read

Version v3.1.7 opens the platform's telemetry to your customers: Metrics Export adds Prometheus-format scrape endpoints on the user API, so a customer can point their own Prometheus or Grafana at the panel and chart their instances, managed databases and load balancers with the API token they already have. The second headline is the Import Doctor: virtual machines imported from Hyper-V, VMware or other platforms - which typically blue-screen on first boot because their disk drivers were never armed for our hardware - are now detected automatically and repaired with one click. Around them, a Fix Monitoring action repairs broken telemetry on managed services in one step, Kubernetes plan rotation picked up the fixes from its first weeks in the field, and the user panel's list pages were made consistent end to end.

Metrics Export

  • [Feature] Prometheus Scrape Endpoints - GET /api/metrics returns every instance, managed database and load balancer the account owns as one Prometheus text exposition, with per-resource endpoints alongside it for narrowly scoped scrape jobs. Authentication is the same bearer token used everywhere else on the user API - there is no separate credential to mint. Metric families are stable and customer-facing (hv_instance_*, hv_db_*, hv_lb_*), with resource id and name labels and no internal topology in the output. Instances include Kubernetes worker nodes and VPN gateway backing VMs.
  • [Feature] Built for Scrapers, Not Browsers - The endpoints answer HTTP 200 always: a backing store being unreachable surfaces as hv_resource_up 0 on the affected resources, never a 5xx that turns into an error storm in the customer's Prometheus. Scrapes are rate limited per token at 30 per minute, the account-wide endpoint caches its rendered body for 10 seconds to absorb multi-target fan-out, and the whole call runs under a hard time budget so one slow backend cannot stall the response. Subuser tokens see exactly the resources their team permissions grant, nothing more.
  • [Feature] Documentation and Examples - The user API documentation gained a Metrics Export section with the full endpoint table and a ready-to-paste scrape_config snippet, and a feature guide ships with this release. The pipeline was validated end to end against a real Prometheus and Grafana stack before release.

Imported Virtual Machines

  • [Feature] Import Doctor - Customers who convert a VHDX or VMDK, write it over their instance disk and boot typically hit INACCESSIBLE_BOOT_DEVICE: installing the virtio drivers inside Hyper-V or VMware stages them but never arms them for boot, because that hardware never presents a virtio disk. The platform now fingerprints every instance disk after each cold start, and when a foreign operating system appears, the instance page shows a banner on both the user and admin panels. One click repairs it offline: the staged virtio storage drivers are armed directly in the guest's registry, stale UEFI boot entries are reset, and the firmware type is matched to what the disk actually uses. Guests where the full repair is not possible fall back to compatible SATA and e1000 emulation so they boot regardless. Detection is automatic; repair only ever runs when someone asks for it, with the guest shut off.
  • [Fix] UEFI Firmware on Debian Hypervisors - UEFI guest definitions probed only the Red Hat and legacy Ubuntu OVMF firmware layouts, so on Debian hypervisors (and newer Ubuntu, which drops the legacy names too) every UEFI define failed with a missing-file error until symlinks were made by hand. Firmware is now resolved from ordered CODE/VARS pairs covering all three layouts, requiring both halves of a pair so a partial install can never mix incompatible images.

Monitoring and Managed Databases

  • [Feature] Fix Monitoring - Managed databases and load balancers gained a Fix Monitoring action that repairs the telemetry pipeline in one step: the metrics agent configuration is re-rendered and re-installed, the metrics exporter login and its grants are re-created, and the agent is restarted. It is available on the metrics tab in both panels - once per resource per 24 hours on the user side, unthrottled for admins - and a failed attempt never consumes the user's daily budget. A new monitoring:fix console command batches the same repair across every active managed service, for fleet-wide rollout after an endpoint or credential change.
  • [Fix] Customer-Created Read-Only Roles - On PostgreSQL, the customer admin account could create a role but not grant it anything useful: granting membership in the monitoring and read-all roles failed with a permission error, so a read-only or Grafana login ended up empty. On MySQL and MariaDB the grant capability had been removed outright. The admin account now carries the delegation rights it needs on both engines - scoped so the hardening that removed superuser-level capabilities stays in force - and the Fix Monitoring action applies the same correction to existing databases.
  • [Fix] Honest Instance Memory Graphs - Instance memory usage was derived from the guest's free-memory figure, which page cache drains toward zero on any warm Linux guest, so memory graphs crept toward 100% regardless of real pressure. The calculation now uses the guest's available-memory figure, which counts reclaimable cache as free - the same number free -m shows in its available column. Guests with older drivers keep the previous behaviour.

Kubernetes

  • [Feature] Cluster API Endpoint by Name - New public clusters serve their API endpoint by their System DNS name (k8s-<name>-cp.<your-domain>) instead of a bare IP: the name is baked into the API server certificate at bootstrap, the cluster page shows it, and generated kubeconfigs use it. The name is pinned at creation and never rewritten, since it lives inside a signed certificate distributed to customers. Existing clusters are untouched and keep working by IP.
  • [Feature] Webhook Subscriptions API - The user API gained full management of webhook subscriptions (/api/webhook-subscriptions): create, list, update, delete, and per-subscription delivery history. Deliveries are HMAC-signed with a per-subscription secret, endpoints must be HTTPS, and subscriptions can filter to a single cluster. Kubernetes lifecycle events are the first event source.
  • [Improvement] Plan Changes Lead Into Rotation - After changing a pool's instance plan, nothing pointed at the rotation that actually resizes the workers, so pools sat reporting drift until someone found the unlabeled icon. The edit form now says what saving will and will not do, a plan-changing save flows directly into the rotation dialog (including the separate downsize confirmation where it applies), and the drift badge itself became a clickable "Rotate now" action. Declining the dialog simply leaves the badge as the reminder.
  • [Fix] Rotation Waves Apply Labels and Taints - Workers added by a plan rotation joined without their pool's labels and taints and showed no role in kubectl get nodes. Each rotation wave now applies the pool's labels and taints after the new workers are ready and before the old ones drain - exactly the moment draining reschedules pods onto them - and the role label appears promptly instead of up to a day later.
  • [Fix] Rotation Dispatch Self-Heals - A rotation is launched as a detached background process, and that launch can die silently. A rotation still pending after three minutes is now redispatched once by the periodic janitor - rotations are resume-safe by design, and a healthy run is never touched - and every detached launch now leaves a forensic log so a failed spawn is diagnosable rather than a mystery.
  • [Fix] Pool Deletion Cannot Strand Workers - Deleting a node pool could remove the pool record while its workers were still protected by scale-down guards, leaving live worker VMs attached to a deleted pool where no cleanup process could reach them. Pool deletion now marks every member for collection with those guards bypassed - they exist to protect a pool that is staying - and refuses to remove the pool record if any live member could not be marked. The cleanup and drain paths also learned to resolve a deleted pool's configuration, so members marked before the deletion still drain correctly.
  • [Fix] Placement on Heavily Committed Nodes - A node whose allocated memory exceeded its physical total crashed the placement query for its whole group with a database range error, aborting scale-ups that had healthy capacity elsewhere. Placement and the rotation capacity gate now both measure headroom against the node's effective memory ceiling - the operator-set overcommit limit where one is configured, physical memory otherwise - so a sanctioned overcommit node counts its real headroom and an over-allocated one simply sorts last instead of taking the group down. The placement query also gained the standard maintenance, lock and deployment gates it was missing.
  • [Fix] Certificate Auto-Renewal Now Scheduled - The cluster certificate renewal service shipped fully built but nothing ever ran it. It now runs daily, renewing certificates inside a 30-day window ahead of expiry.

Networking

  • [Fix] Private DNS on Dual-Homed Instances - On instances with both a public IP and a VPC interface, the first lookup of a private DNS name could stall for seconds: the VPC link shipped without its search domains, so the resolver had no reason to route private-zone queries to the VPC resolver and raced it against the public nameserver. The deploy payload now carries the VPC's DNS zones as search domains. Already-deployed guests pick this up on their next network configuration rebuild.
  • [Fix] Private DNS on VPC-Only Instances - VPC-only instances could lose private-zone resolution entirely: a fallback public nameserver on the same link could permanently win the resolver's affinity, and the VPC's own resolver was not even running until the VPC had at least one zone. The gateway resolver now always runs - from VPC creation, zones or not - and VPC guests use it exclusively on that link, with public resolution forwarded upstream through it.

User Panel

  • [Improvement] Consistent Filtering Everywhere - The service list pages were brought up to one standard: URL-shareable filter state, status and facet filters, and debounced server-side search. Managed database filters that the server always supported are now in the UI; VPN gateway and scaling group search actually filters instead of doing nothing; volume search no longer breaks the table; pagination keeps your filters instead of dropping them. Columns that were already on the wire but never rendered are now shown - private IPs for databases and VPN gateways, location and plan for instances, distribution and region for images.
  • [Improvement] Capacity Errors Reworded for Customers - When a deploy or scale-up fails for a capacity reason - no memory, storage or IP headroom where the resource was requested - the customer now sees a clear "contact support" message instead of raw infrastructure wording that named things they cannot see or fix. Admins are emailed the precise original error (throttled per distinct error), see it unchanged in the admin panel, and it is always logged.
  • [Improvement] AI Assistant Out of Beta - The AI assistant has run long enough in production to drop the beta label. The badge and the settings-page warning are gone; nothing about its configuration changes.

Billing and WHMCS

  • [Feature] Create-Time Add-Ons - External billing provisioning now accepts additional IPv4 addresses (up to 20) and additional disk (up to 5000 GB) at instance creation. The extra disk is placed on the same storage pool as the primary and the capacity check covers the combined footprint. The WHMCS module exposes both as configurable options; they apply at creation only, by design.
  • [Improvement] Hardened WHMCS Module - The module now keeps its own service-to-instance link table (auto-migrated, self-healing from the custom fields it also auto-creates), so suspend, terminate and upgrade no longer depend on an admin having manually created a custom field. Re-running CreateAccount on a linked service refuses to mint a duplicate instance, and terminating a service whose instance is already gone converges cleanly instead of failing forever. The client-area overview page was redesigned to match the panel, with live data, copyable IPs, and one-click SSO into the panel.
  • [Fix] Server Host Resolution - A WHMCS server saved with only the IP Address field filled (Hostname blank) could never load plans or hypervisor groups into product configuration. All module variants now resolve hostname-or-IP consistently, honour the configured port, tolerate a pasted URL, and fail with a named message when both fields are empty.
  • [Improvement] TLS Posture Made Explicit - The billing API client does not verify the master's TLS certificate, because masters are routinely addressed by IP or carry self-signed certificates and verification would break provisioning on those installs. This trade-off is now documented rather than implicit: where possible, point WHMCS at a hostname with a publicly trusted certificate.

Platform and Admin

  • [Fix] Route Debt Sweep - An inventory pass over every registered route closed a set of long-standing gaps: the kubeconfig acknowledgement button now works, users can delete their object storage access keys (the endpoint existed but was never routed), several links that led to pages that do not exist now redirect to the real pages, and two admin route groups gated on permission slugs that could never be granted are now grantable. Dead controllers, models and page stubs were removed.
  • [Fix] Volume Attach Validation - Attaching a volume now verifies the target instance belongs to the same account and derives the instance's hypervisor group correctly (the previous check read a field instances do not have, and failed for every attach). Cluster-managed workers are rejected as attach targets, and subusers now see their account's volumes and instances on the volume pages.
  • [Fix] Web SSH Stability - Web SSH sessions on remote or CDN-fronted masters appeared to disconnect frequently: overlapping readiness polls could each open a connection against a one-time token, and a rejected duplicate would paint a disconnect overlay over the live terminal. Polling is now single-flight and a stale socket can never steal the display from a live one.
  • [Fix] Encrypted Secret Storage - Columns storing encrypted provider secrets were widened; a secret over 23 characters could previously be truncated at rest.
  • [Improvement] PowerDNS Onboarding - Adding a System DNS domain whose zone already exists in PowerDNS now adopts the zone instead of failing, the delegation checker installs its DNS tooling where missing, and a zone error clears automatically once resolved. The wizard's guidance on public-suffix domains was corrected.

Upgrade notes

  • Two migrations run on upgrade: the imported-OS state column and the encrypted secret column widening.
  • Deploy the master before the hypervisor agents for the VPC DNS fixes; each side tolerates the other being old. The agent release carries the memory metric fix, the Debian UEFI firmware resolution and the Import Doctor tooling - the agent installs its guest-inspection packages during provisioning or update.
  • The read-only role grant fix applies to newly provisioned databases automatically. For existing databases, run php artisan monitoring:fix --type=db once (or use the Fix Monitoring button per database) to roll it out.
  • Metrics Export needs no setup: it uses the metrics backends you already configured per hypervisor group. See the new Metrics Export guide for the customer-facing details and a scrape configuration example.

Stable Release Version v3.1.6

· 12 min read

Version v3.1.6 is a large release across Kubernetes, networking and databases. The headline is node-pool plan rotation: changing a pool's instance plan used to be rejected outright while the pool had live workers, so the only way through was to scale to zero and back, a full capacity outage for that pool. Now the plan change is accepted and the platform rotates the workers for you, provisioning replacements before draining anything. Alongside it, System DNS turns provisioning into something that produces real hostnames rather than bare IP addresses, and managed databases on those names present publicly trusted TLS certificates that renew themselves. Load balancers gain host-based routing and can serve internal VIPs for Kubernetes Services. The release also carries a broad set of reliability fixes across Kubernetes operation locking, reverse DNS, backups and the admin panel.

Kubernetes

  • [Feature] Node-Pool Plan Rotation - A pool's instance plan can now be changed while it has live workers. Saving the change starts a surge rotation: new workers on the new plan are provisioned and confirmed ready first, then old ones are cordoned, drained through an escalating ladder, and destroyed, one wave at a time, so the pool never runs below capacity. Scale-down ordering is preserved across the rotation, so the pool does not scramble which node leaves next. Downsizing to a smaller plan is allowed with an explicit warning rather than blocked - it is your cluster - but it takes a second, separate confirmation naming the consequence, so a single click can never start one.
  • [Feature] Rotation Visibility - The pool list shows a drift badge counting how many workers are still on the previous plan, with a Rotate action to start or resume a rotation. The decision data arrives with the page rather than from a probe request, so opening the page cannot consume the rate limit that governs the rotation endpoint itself.
  • [Feature] Worker Pool Identity - Worker nodes now carry a hypervisor.io/node-pool label, following the same convention as the major managed Kubernetes providers, so kubectl get nodes -L hypervisor.io/node-pool shows which pool each node belongs to. Panel hostnames include the pool name too, so the instance lists in both panels no longer show several pools' worth of identically-shaped names. Renaming a pool heals the label on the next reconciliation. None of this costs anything in the database.
  • [Feature] Internal Load Balancers for Services - A Kubernetes Service of type LoadBalancer annotated <prefix>internal: "true" now gets a private VIP inside the VPC instead of a public IP, and the cloud controller reports that private address back to the cluster. Internal services no longer have to be exposed publicly to be reachable, and no change to the cloud controller was required.
  • [Improvement] One Transport for Long Orchestrations - Rolling upgrades, control-plane upgrades and plan rotations no longer run as queue jobs with a fixed ceiling. They run as detached console processes with a task row, an operation lock and a stale-task reconciler behind them, so a long rotation cannot be killed part-way by a queue timeout.
  • [Fix] Kubernetes Operation Locking - Operation locks now correctly detect an already-held lock across the Redis client the platform ships, so concurrent scale, upgrade and rotation operations on one cluster are properly serialised. Lock acquisition also moved outside the surrounding database transactions, so a rolled-back operation releases its lock immediately instead of waiting for the timeout.

DNS

  • [Feature] System DNS - Every public instance, load balancer, managed database and Kubernetes control plane now receives a hostname automatically. Instances are named from their IP in the familiar cloud style (vm-203-0-113-7.cloud1.example.com); named resources use their own name (db-prod-mysql, lb-frontend, k8s-myapp-cp). Records are created at provision time, follow the resource if its IP changes, and are removed when it is destroyed. Where a reverse zone allows it, a matching PTR record is set so forward and reverse agree, and a value the customer set themselves is never overwritten. DNS problems never block or fail provisioning.
  • [Feature] Delegation Wizard - Base domains are onboarded through a four-step wizard: enter the domain and nameservers, delegate at your registrar, then verify. Verification is real rather than advisory - the panel writes a random record into the zone and confirms both that your registrar delegates the domain and that public resolvers can see that record through the delegation, before the domain publishes anything. A daily re-check moves a domain that loses its delegation to a degraded state, where existing records keep resolving but no new resources are assigned to it, and restores it automatically when the delegation returns. A problem local to the panel, such as its own DNS tooling being unavailable, will never degrade your domains.
  • [Feature] Trusted TLS for Managed Databases - A base domain can hold a Let's Encrypt wildcard certificate covering every name under it, issued over the DNS-01 challenge using the zone the panel already controls. Public managed databases receive it automatically and load it without a restart, so customers can connect with --ssl-mode=VERIFY_IDENTITY or sslmode=verify-full against the public trust store instead of trusting a self-signed certificate. One wildcard per domain keeps issuance well inside Let's Encrypt's rate limits, and renewal runs daily from 30 days before expiry, so a failed attempt has a month of retries behind it rather than being a countdown.
  • [Feature] PowerDNS Deployment Kit - A self-contained kit ships with the release for operators who need authoritative DNS to point System DNS and reverse DNS at. It stands up a primary and any number of secondaries replicated by signed zone transfers, where a new zone propagates to every secondary without touching them. One script drives the whole fleet over SSH from the primary: it prints a readiness table per node and changes nothing until every node passes, then converges each node and verifies replication by querying each one directly, failing loudly if any node is not actually serving the zone.
  • [Improvement] Reverse DNS Zone Forms - The reverse zone forms now show the hostname each automatic PTR format actually produces, built live from the prefix and domain you are typing, instead of four opaque option labels. Zone type is chosen from cards and the zone suffix follows the choice.
  • [Fix] PowerDNS Reverse DNS - Corrected the call into the PowerDNS client when setting or rebuilding a PTR record, and added test coverage over that path. ClouDNS providers were not affected.
  • [Fix] Automatic PTR Format Labels - The second and third automatic PTR format options in the reverse zone form now describe the output they actually produce, and the form previews the resulting hostname for each option. If you use format 2 or 3 on an existing zone, check the preview against what you expect before your next change.

Load Balancers

  • [Feature] Host-Based Routing - Load balancer rules can now match on the request host as well as the path, so one load balancer can serve several hostnames to different backends. The match-type options were consolidated at the same time.
  • [Improvement] Self-Healing Configuration - A load balancer left stranded in the configuring state now recovers on its own instead of needing a manual sync, with the grace period derived from how long the agent-side configuration can legitimately take rather than an arbitrary number.
  • [Fix] HTTPS Redirect and Form Wipes - The HTTPS-redirect option on port 80 now applies correctly, and a real-time update arriving while a load balancer configuration form is open no longer discards what you were typing.

Databases and Backups

  • [Feature] Detached Backup Execution - Managed database backups no longer run inside a single blocking connection to the guest. A backup taking more than about 58 minutes used to be killed mid-upload by a transport timeout that scaled with nothing, and the task showed no movement at all between "running" and completion, so a healthy 13 GB backup was indistinguishable from a hang. Backups now run detached with the panel polling progress, reporting transferred bytes live, and the ceiling is a policy you set rather than an artefact of how the command was run. A genuinely dead guest is now detected in about two minutes.
  • [Feature] Backup Run-History Retention - Backup run records are now pruned on a schedule with a configurable retention period, so the history table stays a useful size on long-lived installs.
  • [Improvement] Honest Backup Schedules - The backup settings page now shows the schedule actually in effect rather than a placeholder that could differ from it, and validates a custom schedule when you save it. New installs default to daily.
  • [Improvement] Faster Failures on Broken Egress - The database agent scripts now fail immediately with a named reason when the guest has no route to the internet, instead of hanging until a timeout. The message points at the VPC NAT gateway, which is the usual cause.
  • [Fix] PostgreSQL Incremental Backups - PostgreSQL incrementals are markers over continuous WAL archiving and upload no object of their own. The completion check now recognises that shape and records them correctly, and a marker whose WAL archiving is not actually running is reported as a failure rather than a success.
  • [Fix] Phantom Restore Keys - A storage key recorded against those marker rows could make restore and retention treat an object that does not exist as downloadable. Marker rows no longer carry one.

Platform and Admin

  • [Improvement] Smaller Update Rollback Snapshots - The snapshot taken before an application update now skips logs, caches, images and existing backups, and skips walking those directories at all rather than listing and discarding them. Snapshots on a busy install were hundreds of megabytes of log files.
  • [Improvement] Quieter Admin Lists - The VPC and load balancer list pages no longer reload on every NAT gateway heartbeat. Updates are batched, so a page that was reloading dozens of times a minute now settles.
  • [Fix] Hypervisor Health Flag - The automatic health flag raised when a node stops reporting now clears again on the node's next successful metrics poll, so a node that recovers from a transient blip returns to the deployment pool by itself. Allow Deployments remains the operator-controlled switch, and the admin page now labels which is which.
  • [Fix] Blank Task Status - Task status could arrive in a form the column could not store, leaving the dashboard blank for tasks that were running. The column has been widened and the accepted values are validated where they arrive.
  • [Fix] Missing Hypervisor Uptime - KVM nodes showed "-" for uptime on the hypervisors list and the dashboard.
  • [Fix] VPC Real-Time Updates - Aligned three broadcast channel names with what the panel subscribes to, so VPC views update live again.
  • [Fix] Missing User Timezone Crashed Instance Metrics - Accounts created without a timezone caused the instance metrics endpoints to fail. Timezone resolution now falls back through the account, the system default and UTC in one place, existing accounts are backfilled, and a setting that nothing had ever written - so subuser invitations always fell back to UTC regardless of the configured default - now reads the correct one.
  • [Fix] Supervisor Workers Pointed at a Missing PHP - The queue worker configuration hardcoded PHP 8.3 while the installers have defaulted to 8.4 for some time, so on a fresh install no queue worker would start, and every application update reapplied the mismatch. The PHP version is now substituted where the configuration is deployed, including the unversioned path that EL hosts use.
  • [Fix] Build Pipeline Hardening - The release build now enforces the compatibility flag when targeting multiple PHP versions, records the flags used with each artifact, and adds a verification step plus a smoke test on the target host.
  • [Fix] Billing and Provisioning Debt - Volume backup charges now apply the account balance policy consistently, user records created through every path carry the timestamps the schema requires, and VNC port allocation can no longer hand the same port to two instances.

Infrastructure Agent

  • [Improvement] PHP Entrypoint Pinning - The agent now selects its PHP interpreter explicitly, preferring 8.4 and falling back to 8.3, verifying the candidate actually carries the extensions it needs before committing to it. The update process repairs already-deployed nodes.
  • [Fix] Proxy Survives an Unlinked Node - The SSH proxy no longer fails to start on a node that has not yet been linked to a master, and linking, relinking or unlinking now takes effect without a restart. Unlinking correctly revokes trust.
  • [Fix] Stale Binaries After an Update - An incremental update could leave the previous version of a running binary in place, because a running executable cannot be overwritten in place. The update now replaces them correctly.
  • [Fix] Missing VNC Port Tolerated - A null or zero VNC port arriving from the master no longer produces invalid guest XML or a firewall error.

Integrations

  • [Improvement] OpenTofu Provider and MCP Server - Both gained the node-pool rotation endpoint, and the MCP tool for updating a node pool gained the instance plan field it was missing, so a plan change and its rotation can be driven from infrastructure-as-code or from an AI agent as well as from the panel.

Upgrade notes

  • Deploy the infrastructure agent before the master. The certificate installation command the master sends reaches a route that only exists in the new agent.
  • Three migrations run on upgrade: the task status column widening, the System DNS schema, and the user timezone backfill.
  • Reverse DNS on PowerDNS providers is worth a quick check after upgrading, and if you use automatic PTR format 2 or 3, confirm the form's preview matches the naming you expect.

Stable Release Version v3.1.5

· 5 min read

Version v3.1.5 closes out the 3.1 line. It adds no new product surface - it makes the surface added in v3.1.0 and v3.1.2 behave consistently. The admin REST API now matches the panel and its own documentation, load balancer routing and Proxmox firewall synchronization are more dependable, alert mail is delivered on its own queue, and every instance operation is fully localized. A platform review ran alongside this work; its outcomes are folded into the items below.

Upgrading is routine. There is nothing to reconfigure and no behavior to relearn.

  • [Improvement] Admin REST API consistency - updating a user no longer requires resending a password, the acting administrator resolves the same way on every endpoint, and account credentials are excluded from responses. Load balancer security group rules are readable from the administrative API.
  • [Improvement] Load balancer routing rules honor catch-all matches, applied after the specific rules and before the default backend, so precedence follows the order the rule list reads.
  • [Improvement] Proxmox security group synchronization converges reliably, so firewall rule changes reach running VMs on every cycle, and an explicit deny takes precedence over a broader allow.
  • [Improvement] Notification and alert mail - certificate lifecycle notices, autoscaler token renewals, the admin income digest, and invoice issuance now use the dedicated notifications queue rather than sharing one with live dashboard traffic.
  • [Improvement] Full localization of instance operations, with a build-time guard so a new action cannot ship without its wording.
  • [Improvement] Quota enforcement is identical from the panel and the API, and quota messages distinguish being at your limit from a brief collision with your own concurrent request.
  • [Security] Tighter tenant boundaries on object storage keys and volume plans, stricter identity binding for cluster nodes, and payment capture bound to its originating transaction.
  • [Security] Cluster join credentials are cleared on every terminal outcome, with kubernetes:scrub-cloudcfg provided for existing installs.
  • [Operations] Managed MariaDB monitoring, scheduler health naming, and a single pinned PHP runtime for workers and cron.

Since v3.1.0

If you are upgrading from the 3.0 line, the two releases between it and this one carry the feature work:

  • v3.1.0 adds native Proxmox VE support. Point the platform at an existing PVE 8 or 9 node or cluster with a single API token and manage it beside your KVM fleet - deploys, VPC networking, security groups, backups, snapshots, live migration, consoles, HA, Docker, managed databases, load balancers, Kubernetes, and exact per-NIC bandwidth metering. Agentless, over the Proxmox API.
  • v3.1.2 adds OAuth sign-in with Google, Microsoft and GitHub, enforceable admin OIDC single sign-on, private locations with a built-in request-access workflow, a per-node deployment readiness engine surfaced on the dashboard, the hypervisor list and each node's detail page, and a substantially tougher load balancer with static-IP deploys, captured error state, and long-lived TCP session support.

v3.1.5 builds directly on both: the Proxmox firewall and load balancer improvements below apply to the surfaces those releases introduced.

Kubernetes cluster credential hygiene

Cluster join credentials are meant to be short-lived - they exist while a node is joining and are cleared once it settles. That clearing now runs on every terminal outcome, so a node that fails to join is cleaned up exactly like one that succeeds, and it runs from the periodic reconcilers as well as the join callback, covering a node that never reports back.

Because the behavior changed, v3.1.5 ships php artisan kubernetes:scrub-cloudcfg for existing installs. It is a dry run by default and requires an explicit --force to write, only touches clusters that are fully destroyed, holds a settle window on top of that, includes soft-deleted records, and logs everything it does. Operators upgrading from an earlier 3.1.x can run it once; new installs never accumulate this data.

Admin REST API

Several /api/v1 endpoints behaved differently from the panel and from their own documentation. They now agree:

  • Updating a user no longer requires a password. You can change a name, a role, or a quota without sending a credential, and an unchanged email address is accepted.
  • The acting administrator resolves consistently across instance image listing, user update and delete, VPC enable and disable, and backup creation.
  • Account secrets never appear in responses - API credentials and multi-factor state are excluded from every serialized user object, and the published examples match.

The API manifest, the OpenTofu provider, and the MCP server coverage gates remain in lockstep at 889 endpoints, verified in CI.

Tenant boundaries

Object storage access keys resolve strictly within the owning account on every user-facing route, while the administrative surface keeps its intentionally cross-account view. Volume plans are enforced against the region that offers them, so a plan wired to one location cannot be provisioned into another and pricing stays aligned with your catalog. Cluster node identity is derived solely from the signed token a node presents when it joins.

None of this requires a schema change or any action on your part.

Localization

Every user-facing string in the instance operation pipeline is now translatable, including task names and result messages assembled at runtime. A build-time guard reads the accepted actions directly from the service and fails the build if any of them lacks its wording, so the coverage cannot silently regress.

Stable Release Version v3.1.2

· 11 min read

Version v3.1.2 is the trust and hardening release that follows the Proxmox debut in v3.1.0. It modernizes how people get into the panel - social sign-in for users and enforceable OIDC single sign-on for admins - and how you sell capacity, with private locations and a built-in request-access workflow. Operators get a per-node deployment readiness engine and a substantially tougher load balancer. Underneath, this cycle ran two full platform security sweeps plus dedicated audits of the backup system and managed databases, on both the master and the hypervisor agent.

  • [Feature] Sign in with Google, Microsoft, or GitHub - users can register and log in through OAuth, link and unlink providers from their profile, and auto-link to an existing account only when the provider asserts a verified email.
  • [Feature] Admin single sign-on (OIDC) - bind admin logins to your identity provider with strict subject binding and no just-in-time provisioning, optionally enforce SSO for all admin password logins, and keep a time-limited break-glass path for IdP outages.
  • [Feature] Private locations with request access - lock any location to selected accounts. Locked regions stay visible in the catalog with a lock treatment, users request access in one click, and admins approve or deny from a dedicated queue with email notifications both ways.
  • [Feature] Node deployment readiness - every hypervisor now carries a live readiness checklist (agent, storage, network, capacity, deploy gates) surfaced as a dashboard card, a fleet list badge, and a per-node checklist with failure-specific fix hints.
  • [Feature] Load balancers on allocated static IPs, captured error state (full detail for admins, a subtle banner for users), and per-frontend idle timeouts - TCP frontends now default to one-hour timeouts with kernel keepalives, so SSH and database sessions through an LB no longer drop at 50 seconds.
  • [Feature] Proxmox surface expansion - VM snapshots, instance tags, guest-agent IP discovery, and backup file-restore in the user API; node issues, scheduled backup jobs, and live migration in the admin API; and admin edits to resources, topology, boot order, and NICs now push live to running VMs.
  • [Feature] Security group drop rules on KVM - rule actions are honored end to end, so explicit drop rules override broader accepts, matching the Proxmox behavior.
  • [Improvement] Teams - instance password mails go to the account owner with every instance-manage member in CC. Admin task queue gains one-click pruning and clean deletion.
  • [Security] Two platform-wide security sweeps, defense-in-depth guardrails for the AI assistant, a backup-system audit in three phases, and a managed-database hardening batch. Details below.

Sign in with Google, Microsoft, and GitHub

The login and registration pages now offer OAuth sign-in for Google, Microsoft, and GitHub. Each provider is enabled individually in the admin settings with its own client credentials; nothing shows on the login page until a provider is configured and switched on.

The linking rules are deliberately conservative, because OAuth auto-linking is a classic account-takeover vector:

  • An OAuth identity auto-links to an existing account only when the provider asserts the email as verified. Google must present a true email_verified claim, GitHub only ever returns primary-and-verified addresses, and Microsoft sign-ins are validated against the tenant-verified UPN with the known cross-tenant takeover patterns (nOAuth) explicitly rejected.
  • Sign-ups that arrive without a usable verified email go through a complete-profile step instead of silently creating a half-formed account, and the account write is transactional so a double submit cannot orphan a user.
  • Logged-in users manage linked providers from their profile: connect, view, and unlink, with relinking handled safely.

Admin single sign-on (OIDC)

Admin access can now be delegated to your identity provider - Okta, Entra ID, Keycloak, or any OIDC-compliant IdP:

  • Strict binding. An admin's IdP identity binds on sub (subject), never on mutable claims, and there is no just-in-time provisioning - only pre-existing admin accounts can bind. The first bind is forensically logged, and stale identities are deleted rather than left dangling.
  • Enforcement. Once your IdP is verified, you can require SSO for all admin password logins. The enforcement policy carries a lockout interlock so you cannot switch it on in a state that would lock every admin out.
  • Break-glass. For IdP outages, php artisan admin:sso-break-glass opens a time-limited bypass that expires on its own. It is a deliberate, logged, console-only action.
  • Setup UI. A new Authentication settings tab covers both features, including an OIDC discovery test that validates your issuer before anything is enforced. HTTPS is required and JWT verification is always on.

Private locations and request access

Locations (hypervisor groups) can now be restricted per account. The catalog stays honest about what exists:

  • Locked regions render on every create surface - the deploy modal, Cloud Service, self-provisioning, VPC and Kubernetes pickers - with a frosted lock treatment and the region name still visible, instead of vanishing from the catalog.
  • Users hit Request access on a locked location, confirm, and the request lands in a new admin queue with a navigation badge. Admins approve or deny inline; both outcomes notify the user by mail. Access states are tracked per account as available, requested, or locked.
  • Admin user pages gain a Cloud Service tab consolidating the account's location grants, inline approve and deny, and the account's provisioning limits.
  • A per-account cloud provisioning switch cleanly disables self-service provisioning for an account without touching its running services, and the billing-exemption logic was made consistent across every surface that renders a deploy button.

Access enforcement is server-side on every create path, across web, API, queue, and AI-assistant surfaces. The lock UI is presentation; the gate is in the services.

Node deployment readiness

Answering "why is nothing deploying to this node" used to mean reading logs. Now every hypervisor - KVM and Proxmox - carries a readiness engine that evaluates the conditions a deploy actually requires: agent reachability, storage presence and free capacity, subnet availability, deploy flags, maintenance and lock state.

  • The admin dashboard shows a fleet readiness card.
  • The hypervisor list badges each node ready, pending, or blocked.
  • The node detail page renders the full checklist, and every failed check carries a specific fix hint tied to the actual failure, not a generic message.
  • Adding a Proxmox node now runs its first cluster reconcile synchronously, so a freshly linked node reports honest readiness immediately instead of waiting for the next cron pass.

The checks mirror the real deploy gates - a node the checklist calls ready is a node the scheduler will actually use.

Load balancer improvements

  • Static IP deploys. User load balancers can deploy onto allocated static IPs, so an LB's address can be planned, firewalled, and DNS'd before it exists.
  • Error surfacing. LB provisioning and sync failures are captured as a last-error state: admins see the full detail on the LB page, users see a subtle banner that something is being worked on - operational detail stays internal.
  • Long-lived TCP sessions. TCP-mode frontends previously inherited HTTP-tuned 50-second idle timeouts, which silently killed idle SSH, database, and message-queue connections through the LB. TCP frontends and backends now default to one-hour timeouts with kernel TCP keepalives on both sides, websocket tunnels get a matching post-upgrade timeout, and every frontend gains an optional idle timeout field (30 to 86400 seconds) in both the user and admin panels for workloads that need more or less.
  • Kubernetes LB fixes. Service LBs honor the managed-loadbalancer-public-ip annotation, weighted routing-rule backends materialize correctly with collision-free ACL names, port 80 stays plaintext under global SSL mode, and NodePort backends are health-checked over TCP.
  • Plan enforcement. Standalone LB deploys enforce the location's plan-group offering, closing a path where an LB could deploy from a plan the region does not sell.

Proxmox, continued

v3.1.0 shipped the driver; v3.1.2 finishes the surfaces around it:

  • User API: VM snapshots (list, create with optional RAM state, rollback, delete), instance tags, guest-agent IP discovery, and backup file-restore browse and download.
  • Admin API: node issues (list, retry, resolve), PVE scheduled backup jobs, and live migration with precheck. Route binder failures return real 404s instead of leaking existence.
  • Live VM edits. Admin changes to resources, CPU topology, boot order, and NICs push to the running VM where PVE allows it, with CPU flags, secure boot, and TPM handling brought to parity with KVM.

All new endpoints are covered by the API manifest and mirrored in the OpenTofu provider and MCP server coverage gates.

Reliability: Kubernetes, VPN gateways, VPC

  • Kubernetes: worker-pool scale-up crash fixed, long jobs no longer double-execute after 90-second queue redelivery, control-plane and worker plan pickers are scoped to the region's plan groups, node selection prefers the NAT-active hypervisor, and deploys survive recycled-IP ARP staleness and transient agent transport blips. Control-plane LB deploys from the queue were failing on an authentication-context assumption; provisioning gates now evaluate the acting user everywhere.
  • VPN gateways: peer key pairs auto-generate as the UI always promised, and road-warrior clients receive the VPC's private DNS resolver.
  • VPC on KVM: cross-node NAT egress now installs the correct default route on non-active nodes and repairs it in the periodic sync, the VPC bridge joins a firewalld zone so nftables cannot silently reject its traffic, and ICMP redirects are suppressed on VPC veths - closing a class of "works from one node, dead from another" reports.
  • Node provisioning: fresh hypervisors install required CLIs rather than only upgrading existing ones, Debian contrib is enabled across both source layouts for ZFS, and half-merged /usr systems are repaired so kernel modules and ufw work on broken base images.

Managed database hardening

The managed database service went through a dedicated audit. Highlights: six critical backup, restore, and HA defects fixed; incremental backup chain source pinning so a restore can never mix chains; encryption keys moved off process argv; a watchdog that rescues clusters stuck in configuring with init-phase visibility; callback token lifecycle hardening with a localhost guard; PostgreSQL cluster self-heal; replica resync credentials forwarded correctly; and the admin password revealed on the detail pages where operators actually need it.

Backup system audit

A three-phase audit of the backup pipeline shipped on both sides:

  • Master: failure alerting is throttled and queue-routed so it always sends, repeated failures auto-pause a plan instead of burning nightly cycles, prune notifications report what was actually pruned, backup sizes are captured from the agent callback, and remote restores gained a direct download path while a dead restore path was removed.
  • Agent: a credential leak into backup artifacts was stopped, silently truncated backups are now detected and failed, backup and restore state files are no longer world-readable, and qcow restores verify the staged artifact and use tmp-then-rename so a partial download can never replace a disk.

Security sweeps

Two platform-wide sweeps (2026-07-29 and 2026-07-30) ran during this cycle, with every finding remediated before release. The notable classes:

  • Billing integrity: top-up capture is now bound to its originating transaction, closing a credit-fraud path; credit adds are validated; backup debits are atomic.
  • Tenant scoping: SSH sessions, S3 access keys, Kubernetes certificate renewal, and VPC selection are all bound to the owning tenant; state-changing restore moved off GET.
  • Auth: password-reset throttling, no exception reflection to clients, OAuth and email uniqueness guarantees, and the Microsoft cross-tenant (nOAuth) rejections described above.
  • Secrets at rest and in transit: queue payloads carrying secrets are encrypted, failed-job rows are pruned, Kubernetes join credentials no longer travel through cloud-init user data, notification channel secrets are no longer serialized into events, and WireGuard AllowedIPs are validated before any privileged guest execution.
  • AI assistant guardrails (three phases of defense in depth): streaming egress redaction of configured secrets, knowledge-base audience scoping that fails closed, untrusted-data framing around tool output with prompt-injection guards, and redaction of persisted tool calls and audit logs so the assistant's own storage cannot become the leak.
  • Dependencies: dompdf bumped for CVE-2026-56722.

Stable Release Version v3.1.0

· 11 min read

Version v3.1.0 is the Proxmox release. The platform now speaks Proxmox VE natively: point it at any existing PVE 8 or 9 node or cluster with a single API token, and every node is discovered, linked, and managed from the same panel, API, and billing pipeline as your KVM fleet. There is nothing to install inside Proxmox and no agent to maintain - the driver works entirely over the Proxmox REST API, with one optional one-command bootstrap on a node for the features that need host access (VPC WebSSH and exact per-NIC metering). Nothing changes for existing KVM hypervisors, instances, plans, or balances.

  • [Feature] Native Proxmox VE driver - full instance lifecycle on PVE 8.0+ and 9.x: deploy, power, suspend and resume, reinstall, resize, destroy, ISO mounting, cloud-init, GPU passthrough, VM tags, and per-tenant resource pools. Agentless by design; see the section below.
  • [Feature] One-token cluster onboarding - give the panel one reachable member IP and one API token, and it discovers every node in the cluster, links them all, verifies SDN and FRR prerequisites, provisions console credentials, and self-checks the token's privileges. Failures are listed per node instead of aborting the lot.
  • [Feature] VPC networking on Proxmox SDN - VPCs are realized as PVE SDN EVPN zones, vnets, and subnets, with NAT gateways, hot attach and detach of VPC NICs, and automatic reconciliation. Security groups translate to the PVE firewall with accept and drop rules, IP sets, and per-VM rulesets.
  • [Feature] Backups and snapshots - vzdump backup, restore, and delete per instance, optional live-restore, native PVE scheduled backup jobs with full retention control, file-level restore (browse a backup and download individual files), and VM snapshots with optional RAM state, rollback, and delete.
  • [Feature] Live migration - move a running VM between cluster nodes from the panel, with an eligibility precheck, local-disk handling, target storage selection, bandwidth limits, and a dedicated migration network option.
  • [Feature] High availability via the PVE CRM - instances enroll into Proxmox's own HA stack on deploy, with per-group tuning and placement rules; the panel mirrors CRM state instead of running a second watchdog.
  • [Feature] Full guest-service parity - Docker deployments, managed databases (including S3 backups, point-in-time recovery, and replication), load balancers with Let's Encrypt, Kubernetes clusters, WebSSH, SSH key injection, and password resets all work on Proxmox instances through the QEMU guest agent.
  • [Feature] Exact bandwidth metering - per-NIC counters are read from the host, so VPC (east-west) traffic and public traffic are billed separately and never double-counted, with reset-safe accumulation across reboots.
  • [Feature] Node Issues dashboard - operational failures on a Proxmox node (metering, host access, proxy installs, firewall gaps) surface as first-class admin issues that reopen if they recur, instead of hiding in a log file.
  • [Improvement] Fleet-scale collection - metrics, statistics, security-group reconciliation, and HA reconciliation all fan out per node onto a worker pool, so a large cluster is collected in parallel rather than serially.

Native Proxmox VE support

v3.1.0 introduces a second hypervisor backend alongside KVM. A hypervisor group is now either a KVM group or a Proxmox cluster, and both kinds run side by side on one panel with the same instances, plans, billing, and user experience.

The design principle is API-only: the driver talks to the Proxmox REST API and nothing else. Your cluster keeps looking like a normal Proxmox cluster - VMs created by the panel are plain QEMU VMs you can see in the PVE web UI, tasks the panel starts are ordinary PVE tasks (UPIDs the panel tracks and can cancel), and nothing on the node is patched or replaced.

Onboarding an existing cluster

Create a hypervisor group of type Proxmox, paste one member's address and one API token, and the panel calls the cluster status endpoint to discover every node. Each node is linked individually, storage is discovered per node with its real PVE plugin type, an @pve console credential is provisioned for VNC, and SDN plus FRR prerequisites are verified so VPC problems are caught at link time rather than at first deploy. The token itself is self-checked: the panel reports the PVE version and the privileges the token actually holds, and flags the guest-agent privilege gap that PVE 9 introduced, right on the node's detail page.

Supported: Proxmox VE 8.0 and later, including 9.x, standalone nodes and clusters. Group names must match the PVE cluster name so the group is an unambiguous mirror of the datacenter, and groups are kept homogeneous - one group is either KVM or Proxmox, never a mix.

Images and deploys

Operating system images are built once per cluster into a vzdump archive, then every deploy restores that archive into the customer's VM ID. There are no template VMs parked on your nodes and no full-clone storms: what lands is a plain VM, on the storage you chose, with cloud-init applied. Archives are labeled with the upstream image name, can be purged with one command, and image builds are excluded from orphan detection.

Cloud-init works two ways: the default NoCloud seed ISO (full feature set, including multi-user and multi-IP layouts), or PVE's native cloud-init as a per-group opt-in for simple single-user, single-IP instances where a config drive is unnecessary.

Existing VMs on the cluster are not stranded either: the orphan importer can adopt any VM the panel did not create - including raw, ZFS, LVM, and Ceph-backed disks - and attach it to a user as a managed instance.

Storage

All common PVE storage backends are supported and detected with their real capabilities: directory, NFS, LVM, LVM-thin, ZFS, and Ceph RBD, including onboarding an external Ceph cluster directly from the panel. Extra volumes allocate through the PVE storage API with the correct format per backend, hot-plug with IO limits, and resize live (grow-only, as QEMU requires). Capability gates are honest: operations a given backend cannot do, such as snapshots on plain LVM, are refused with a clear message instead of failing halfway.

Networking, VPCs, and security groups

VPCs on a Proxmox group are built on PVE SDN: an EVPN controller and zone per VPC with a dedicated VRF VNI, a vnet and subnet per VPC subnet, and a NAT gateway with SNAT semantics that hold cluster-wide. VPC NICs hot attach and detach on running VMs. A reconciler keeps the SDN objects healthy and converges NAT gateway activation, and VPC create and destroy push eagerly so the fabric follows the panel immediately.

Security groups compile to the PVE firewall: cluster-level groups and IP sets, per-VM rulesets with both accept and drop actions, anti-spoofing IP filters, and firewall flags only on the NICs they belong to. Reconciliation is fingerprint-based, so unchanged nodes are skipped, and a cron backstop repairs drift. The datacenter firewall is enabled safely, preserving management access rather than default-dropping the hosts.

The instance network tab also gains guest IP discovery on Proxmox: addresses are read live from the QEMU guest agent, filtered of loopback and link-local noise, for both admins and users.

Consoles and WebSSH

Three console paths ship for Proxmox instances:

  • VNC (noVNC) through a PVE login ticket with single-use, server-side session tokens. PVE credentials and tickets never reach the browser.
  • Serial terminal (xterm.js over termproxy), tenant-scoped.
  • WebSSH - a real SSH shell in the browser. Public-IP guests are dialed directly. VPC-only guests are reached through proxmox-ssh-proxy, a small node-side service that enters the VPC's VRF on the host, because the guest's route only exists there.

The proxy is the one place host access is needed, and it bootstraps with one command on one node: the panel mints a short-lived, single-use install command, the script fetches the binary from your panel (never a third-party artifact), installs the service, opens the master's port in the PVE firewall using the egress address the node itself observed, and the panel keeps every node's proxy current automatically from then on. Install state is verified against what is actually on the node, so a half-finished install shows up as fixable instead of hiding.

Backups, scheduled jobs, and file restore

Per-instance backups run through vzdump with completion tracked to the panel's backup queue, restore (optionally live-restore, so the VM boots while data streams back), redirect-to-storage on restore, and delete. Cluster-level scheduled backup jobs are managed from the panel with the full keep-last, keep-hourly, keep-daily, keep-weekly, keep-monthly, and keep-yearly retention set, targeting specific VMs, a resource pool, or the whole cluster. With Proxmox Backup Server storage, file-level restore lets admins and users browse a backup's filesystem and download individual files without restoring the VM.

Snapshots and Forge

VM snapshots are first-class on capable storage: create with or without RAM state, list, rollback, and delete, from both admin and user instance pages. Forge (the checkpoint-try-rollback workflow) rides the same mechanism with a reserved snapshot slot, protected from name collisions with user snapshots.

Live migration

Admins can live-migrate a Proxmox instance between cluster nodes with a precheck that refuses ineligible targets and local hardware blockers, handles local disks, and passes through target storage, bandwidth limit, migration network, and online mode. On PVE 9, conntrack-state migration is available as an opt-in. Migration success updates the instance's node assignment and flushes routing caches; a failed migration leaves the instance exactly where it was.

High availability

Proxmox groups delegate HA to the PVE cluster resource manager - the layer that actually owns fencing and recovery - instead of duplicating it. Instances enroll on deploy and unenroll on destroy, groups expose HA tuning (and node-affinity placement, using HA groups on PVE 8 and node-affinity rules on PVE 9), and a reconciler mirrors CRM state back into the panel so the instance page shows the real HA state. The KVM HA watchdog is untouched, and a Proxmox-side failure can never stall KVM monitoring.

Guest services: Docker, databases, load balancers, Kubernetes

Everything that previously required the slave agent's SSH path now runs over the QEMU guest agent on Proxmox instances:

  • Docker: deploy, control, status, and logs.
  • Managed databases: the full suite - configure, restart, password reset, backup tooling, full and incremental backups to S3, point-in-time recovery, restore, upgrade, replication, and batched health checks.
  • Load balancers: configuration, Let's Encrypt issue and renew, and telemetry setup.
  • Kubernetes: kubectl execution, control-plane certificate upload and renewal, and admin kubeconfig reissue, routed through the Proxmox adapter for Proxmox-hosted clusters.
  • Instance basics: SSH key inject and remove with port detection, password resets, and WireGuard configuration sync.

Large payloads are delivered safely: small files write through the agent directly, larger ones are pulled by the guest from a single-use, checksum-verified URL.

Metering, billing, and scale

Metrics and statistics feed the same billing pipeline as KVM, with the units and counters aligned so plan enforcement and charts behave identically. Two things are worth calling out:

  • Exact per-NIC bandwidth. The PVE API only exposes aggregate VM counters, which cannot split public from VPC traffic. v3.1.0 reads per-NIC tap counters from the host instead, so VPC traffic and public traffic are metered separately and never double-billed, with reset-safe accumulation that keeps mid-day reboots billable.
  • Parallel collection. Metrics, instance statistics, HA reconciliation, and security-group reconciliation fan out per node and per group onto a dedicated worker pool. Collection time scales with your largest node, not your node count.

Operational visibility

A new Node Issues dashboard (admin, under Proxmox) surfaces operational failures on a node as structured issues: what failed, on which node, how many times, first and last seen. Issues reopen automatically if the condition recurs after being marked solved, and each issue carries a retry action. Long-running PVE tasks are tracked by UPID and can be cancelled from the panel, and a stuck panel task cancels its PVE counterpart when marked failed.

Security hardening in this release

The Proxmox driver went through dedicated security review during the cycle. Notable fixes shipping in v3.1.0: a WireGuard configuration path that could execute as root in the guest was closed, WebSSH sessions no longer expose PVE authentication cookies, console ACLs are revoked rather than accruing, ISO URL fetching is SSRF-hardened, console routes are tenant-scoped with negative-authorization tests, and backup restore and file-restore endpoints verify the backup belongs to the instance before acting.

Stable Release Version v3.0.0

· 6 min read

Version v3.0.0 is a foundation release. The platform - both the master and the hypervisor agent - now runs on Laravel 13 on PHP 8.3, and this build adds a broad round of precision and edge-case hardening across the systems that run quietly in the background. It also ships two ways to drive the platform programmatically: the Hypervisor.io OpenTofu / Terraform provider, so your instances, networks, and clusters can be managed as code, and a remote Model Context Protocol (MCP) server, so AI agents can operate the same account conversationally. Nothing changes for your existing instances, plans, balances, or configuration.

  • [Feature] OpenTofu / Terraform provider - manage your account as Infrastructure as Code. The iaas provider (now released at v0.2.1, source hypervisor-io/iaas) exposes 56 resources and data sources - instances, VPCs and subnets, Kubernetes clusters, managed databases, load balancers, storage, DNS, VPN, S3, and more - over the user REST API, with full CRUD, import, and reference-driven dependencies. Authentication is your existing IP-locked API token, and both users and admins can use it. See the section below.
  • [Feature] MCP server - hand your account to AI agents. A remote, stateless Streamable HTTP Model Context Protocol server exposes the platform as 364 tools (302 user + 62 curated admin) over the same REST API, with Bearer API-token pass-through auth, confirm-gated destructive operations, idempotency keys, and async convergence. Any MCP client can drive it. See the section below.
  • [Platform] Laravel 13 across the stack - master and the hypervisor agent both move to Laravel 13 on PHP 8.3, on stable, security-audited dependencies. A pure runtime modernization; your data and settings are untouched.
  • [Improvement] Billing precision - hourly billing is refined for long-running resources, unusual calendar edge cases, storage metering cadence, and exact fractional-credit accounting.
  • [Improvement] Backup & restore robustness - incremental chains, in-place volume restore, and retention pruning were hardened for edge cases across Ceph and block backends.
  • [Improvement] Migration & networking polish - cold and live migration edge cases, reverse DNS for uncommon IPv6 forms, upload/download rate-limit symmetry, and quota handling under concurrency were all tightened.
  • [Improvement] Metrics accuracy - the usage pipeline that feeds billing now handles counter-reset edge cases more precisely.

Infrastructure as Code (OpenTofu / Terraform)

v3.0.0 introduces the Hypervisor.io OpenTofu / Terraform provider - the whole platform, declared in HCL and converged with tofu apply.

How the OpenTofu provider fits together

You write the resources you want; the iaas provider translates them into calls against the user REST API using your existing IP-locked API token, and OpenTofu tracks the state. It covers 56 resources and data sources - instances, VPCs and subnets, security groups, Kubernetes clusters and node pools, managed databases, load balancers, storage volumes, DNS, VPN gateways, S3 buckets, projects, and the catalog data sources you reference for plans, images, and regions. Every resource supports full create / read / update / delete, tofu import <addr> <uuid> to adopt existing infrastructure, and normal Terraform references so one resource can depend on another.

Both users and admins can use it: a user manages their own resources with a user token; an operator uses an admin-scoped token for the broader surface. Because the token is validated against the IP it was registered with, run tofu from a stable egress IP (a CI runner, bastion, or workstation).

The provider is released at v0.2.1. Declare it with source hypervisor-io/iaas and let OpenTofu fetch it:

terraform {
required_providers {
iaas = {
source = "hypervisor-io/iaas"
}
}
}
  • Provider and full documentation: the repository is github.com/hypervisor-io/terraform-provider-iaas and the reference docs walk through getting started, the resource catalog, and common patterns.
  • Publishing to the public OpenTofu / Terraform registry is being finalized; that hypervisor-io/iaas source is the install path once it lands, and you can build the provider from the repository in the meantime.

AI agents (Model Context Protocol server)

v3.0.0 also ships the Hypervisor.io MCP server - the same platform, exposed to AI agents. Where the OpenTofu provider is the declarative path, the MCP server is the conversational one: point any MCP client at it and an agent can create instances, assign VPCs, and manage the rest of your infrastructure in natural language.

The API is the source of truth; the OpenTofu provider and the MCP server are two consumers over one Go client

It is a remote, stateless Streamable HTTP server, so there is nothing to install locally - connect over HTTP and authenticate with a Bearer API token that is passed straight through to the platform, IP-locked exactly like the provider's. It exposes 364 tools - 302 user tools plus 62 curated admin tools - covering instances, VPCs and subnets, Kubernetes, managed databases, load balancers, storage, DNS, VPN, S3, and the catalog lookups agents need to reference plans, images, and regions.

Because agents act on their own, the server is built to be safe by construction:

  • Confirm-gated destructive operations - deletes and other irreversible actions require an explicit confirmation step before they run.
  • Idempotency keys - a retried call does not double-create, so a flaky connection cannot spawn duplicate instances.
  • Async convergence - long-running operations return a task the agent can poll to completion instead of blocking.
  • A curated admin allowlist - the 62 admin tools are a deliberately safe subset; billing, user deletion, and hypervisor destruction are not exposed.

The server is backed by the same tested Go client as the OpenTofu provider, so both consumers reach the API through one audited code path.

The user / admin REST API is the single source of truth for all three surfaces - the API itself, the OpenTofu provider, and the MCP server - and a manifest-driven CI gate keeps them in lockstep, so no endpoint can ship without matching provider and MCP coverage.

Beta Release Version v2.2.6

· 6 min read

Version v2.2.6 makes local storage a first-class, pluggable layer. Until now a local instance disk meant a qcow2 file. This release adds four block-device backends - LVM thin, LVM thick, ZFS thin, and ZFS thick - that slot in alongside qcow2, Ceph, and NFS with the full lifecycle: deploy, resize, snapshots, incremental backups, restores, image creation, and live migration. Around them sit three operator-facing safety features: a guided storage wizard that probes the hypervisor before anything is saved, a collapse watchdog that suspends only the at-risk instances before a thin pool can corrupt, and a live Storage Health dashboard. The Virtualizor importer now lands migrated VMs straight onto LVM/ZFS with no qcow2 conversion step, and the admin panel picked up a round of polish along the way. This is a beta build - feature-complete for the storage work and safe to stage, but not yet promoted to stable.

  • [Feature] LVM and ZFS Storage Backends - Hypervisor storage now supports four new local pool types alongside qcow2 files, Ceph, and NFS: LVM thin, LVM thick, ZFS thin, and ZFS thick. Instances on these pools run on raw block devices for near-native disk performance, with the thin variants providing over-provisioning and the ZFS variants bringing checksumming and inline compression. Plans can target a specific backend, so an operator can sell "NVMe block storage" and "standard storage" tiers from the same hypervisor.
  • [Feature] Guided Storage Wizard - Adding a storage pool is now a five-step wizard that probes the hypervisor live before anything is saved: it verifies the tooling is installed, the volume group or pool exists, and reports real capacity, so a mistyped pool name can no longer create a broken storage row. LVM and ZFS appear in the picker with a Beta tag for this release.
  • [Feature] Storage Collapse Watchdog - Every hypervisor now watches its pools every 30 seconds. When a thin pool or ZFS pool approaches exhaustion, the platform suspends only the affected instances before their disks corrupt, alerts every administrator by mail, and surfaces the event live in the Storage Health dashboard. Recovery is operator-controlled: fix the pool, then resume the instances from where they stopped.
  • [Feature] Storage Health Dashboard - A per-pool health view in the admin panel shows live data and metadata usage, ZFS pool state and fragmentation, and warn/critical thresholds, fed by the same probes the watchdog uses, surfaced inline on the storage list.
  • [Feature] Native LVM/ZFS Virtualizor Import - The Virtualizor importer no longer forces every disk through a qcow2 conversion. When the source VM lives on LVM or ZFS, the slave probes the pool and the importer lands the instance natively on the matching backend, preserving the storage tuple and skipping the conversion pass entirely.
  • [Improvement] Incremental Backups Everywhere - The QEMU dirty-bitmap incremental backup pipeline now covers all seven backends, including raw block devices, writing the same qcow2 backup chains to your existing backup storage. Existing qcow2 and Ceph backup behavior is unchanged.
  • [Improvement] Live Migration for Block Pools - Live migration works between LVM/ZFS pools across hypervisors, pre-creating the destination volume and block-copying storage as part of the move.
  • [Improvement] Admin Panel Polish - The VPN Gateway, Managed Database, and Load Balancer plan create/edit pages were rebuilt on the shared sectioned-card layout used by Instance Plans (and a latent dropdown bug on those pages was fixed), more list pages gained the standard collapsible filter panel, the in-browser SSH terminal now stays connected for as long as the tab is open, and the frontend now builds through Vite.

LVM and ZFS Storage Backends

Until now local instance disks meant qcow2 files. v2.2.6 adds four block-device backends - LVM thin, LVM thick, ZFS thin, and ZFS thick - that slot in alongside the existing qcow2, Ceph, and NFS options with the full feature set: deploy, resize, snapshots (thin pools and ZFS), incremental backups, restores, image creation, and live migration. Because plans can target a specific backend, an operator can offer differentiated storage tiers from a single hypervisor.

Pools are added through the new guided wizard, which talks to the hypervisor before saving: it checks that lvm2 or zfsutils are installed (both are installed when you update the hypervisor from the admin panel), that the volume group, thin pool, or zpool actually exists, and shows live capacity. Existing storage rows, plans, and instances are untouched - qcow2, Ceph, and NFS behave exactly as on 2.2.5, and nothing is migrated or rewritten on upgrade.

The Collapse Watchdog

Thin provisioning's dark side is pool exhaustion: when an over-provisioned pool fills up, every VM on it can corrupt at once. v2.2.6 ships a two-layer defense. First, LVM thin pools are configured to auto-extend at 80% via dmeventd. If that cannot help - the volume group is out of space, or a ZFS pool faults - a watchdog on each hypervisor catches the threshold crossing within 30 seconds, pauses only the instances on the affected pool, mails every administrator, and raises a live alert in the Storage Health dashboard. Paused instances keep their RAM and disk state; once the operator extends the pool or replaces the failed device, the instances resume from where they stopped. Degraded ZFS mirrors raise an early warning while the instances keep running.

Storage Health Dashboard

The Storage Health view in the admin panel shows per-pool data and metadata usage, ZFS pool state and fragmentation, and warn/critical thresholds - fed by the same probes the watchdog uses and broadcast live over WebSocket. In this release the health values are surfaced inline on the storage list, so pool state lives in one place.

Native LVM/ZFS Virtualizor Import

Migrating off Virtualizor no longer means rewriting every disk to qcow2. When a source VM lives on LVM or ZFS, the slave probes the pool and the importer lands the instance natively on the matching backend, preserving the storage tuple and skipping the conversion pass entirely. qcow2 sources continue to import as qcow2.

Known Beta Notes

  • LVM and ZFS appear with a Beta tag in the storage wizard; exercise them on a test hypervisor before offering them to customers.
  • Pool auto-extend depends on free space in the underlying volume group or zpool - the watchdog is the safety net, not a substitute for capacity planning.
  • Found a problem? Report it referencing v2.2.6-beta along with the hypervisor's storage type and the relevant Storage Health alert.

Release Version v2.2.5

· 10 min read

Version v2.2.5 turns the platform into a place you can ship code, not just run servers. The headline is Git Push-to-Deploy: customers connect GitHub, GitLab, Bitbucket, or Gitea, point the platform at a repository, and every push is built and rolled out to their instance with zero downtime - no Dockerfile required, thanks to automatic build packs. Pull requests get their own disposable preview environments, complete with status comments posted back to the PR. Around the headline, the release adds live master updates so operators can upgrade the platform from a banner in the admin panel and watch it happen, unicast VXLAN so VPC networking works in datacenters without multicast, per-interface VPC speed controls, a new bandwidth overage billing option that charges for extra traffic instead of cutting customers off, and a platform-wide security hardening pass.

  • [Feature] Git Push-to-Deploy - Deploy applications directly from a Git repository to an instance through the Docker Manager. Pick a repository and branch, choose how it builds, and the platform clones, builds, and runs it. From then on, every push to the branch deploys automatically.
  • [Feature] Four Git Providers - Connect GitHub through a guided GitHub App flow with repository and branch browsing built into the panel, or connect GitLab, Bitbucket, and Gitea. Private repositories are supported through access tokens or SSH deploy keys, including custom SSH ports and users for self-hosted servers.
  • [Feature] Automatic Builds with Build Packs - Applications build with Nixpacks or Railpack, which detect the language and framework and produce a container image with no Dockerfile in the repository. A plain Dockerfile and a static-site build pack are also available. Builds run inside an isolated helper container, so no build tooling is ever installed onto the instance itself.
  • [Feature] Zero-Downtime Rolling Deploys - A new deployment starts the new container, waits for it to come up healthy, switches traffic over, and only then removes the old one. Traffic is routed through a managed reverse proxy with automatic HTTPS certificates.
  • [Feature] Pull-Request Preview Environments - Every pull request can get its own live preview deployment on its own URL. The preview is updated on each new commit, the deployment status is posted back to the pull request as a comment, and the environment is torn down when the PR closes. A previews list in the panel shows what is running and allows manual teardown.
  • [Feature] Dedicated Build Hosts - Builds can be offloaded from the application instance to a separate build host. Each user gets their own deploy keypair, and the panel includes a guided onboarding flow with a connection test before the host is used.
  • [Feature] Deploy Controls and Logs - Webhook deploys are signature-verified, a deployment can be skipped by putting [skip ci] or [skip cd] in the commit message, and full build and deploy logs are available in the panel with credentials automatically redacted.
  • [Feature] Live Master Updates - When a new platform version is available, an update banner appears in the admin panel. One click starts the upgrade, and a live console streams the update progress in real time until it completes. The dashboard now also carries a version pill showing the running release.
  • [Feature] Unicast VXLAN for VPC Networking - Each hypervisor group can now run its VPC overlay in unicast mode instead of multicast. Unicast mode works in datacenters and on networks where multicast is not available, which removes the most common blocker to enabling VPC networking. Existing groups keep multicast by default.
  • [Feature] VPC Interface Speed Controls - Operators can set default inbound and outbound speed limits for VPC network interfaces per hypervisor group, and override them per instance plan. Customers see the effective limits on their instance's network details.
  • [Feature] Bandwidth Overage Billing - Instance, managed database, load balancer, and VPN gateway plans gain a Charge Overage option. When a customer exhausts the plan's bandwidth allowance, traffic keeps flowing and the extra usage is billed per gigabyte at the rate set on the plan, charged through the hourly billing cycle. The existing cut-off behavior remains available for plans that prefer it.
  • [Feature] Expanded Instance Charts - Instance monitoring gains disk I/O, disk usage, and network packet and error charts alongside the existing CPU, memory, and bandwidth graphs.

Release Version v2.2.4

· 11 min read

Version v2.2.4 is our first stable, generally available release. After a long beta cycle that delivered VPC networking, load balancers, object storage, managed databases, and managed Kubernetes, the platform now graduates to stable. This release is headlined by Integrated Payments and Billing, a complete, built-in way to take real money from customers without bolting on an external billing system. Customers top up their account credit directly with a card or wallet, receive automated tax-compliant invoices by email, and can request refunds, all from the same panel they already use to run their infrastructure. Operators get connectable payment gateways, configurable tax rules, promotions, know-your-customer verification, and a revenue dashboard with exportable reports and a daily income summary email. The release also brings self-service plan changes with automatic proration, a major expansion of the AI Assistant so customers and operators can now create and manage resources in plain language, and a comprehensive Admin REST API for automation and integration.

  • [Feature] Integrated Payments and Billing - A complete built-in billing system. Customers add funds to their account balance and that credit pays for hourly Cloud Service usage. No external billing platform required. Works alongside the existing WHMCS, Blesta, and HostBill integrations for operators who prefer them.
  • [Feature] Payment Gateways - Connect Stripe, Razorpay, or PayPal and start accepting payments in minutes. Multiple gateways can run side by side, and customers pick their preferred option at checkout. Payment confirmations are verified securely before any credit is granted.
  • [Feature] Automated Invoices - Every successful top-up generates a sequential, tax-compliant invoice. The invoice is emailed to the customer with a PDF attachment and is always available to view or download from the billing area.
  • [Feature] Tax Rules and Tax Profiles - Operators define tax rules by region. Customers fill in a tax profile with their business and tax-identification details, with built-in validation for European VAT identification numbers so business-to-business transactions are handled correctly.
  • [Feature] Refunds - Customers can request refunds against eligible transactions, and operators approve and process them from the admin billing area. Credit and revenue records stay consistent throughout.
  • [Feature] Promotions - Issue promotional codes that grant bonus credit or a discount at top-up time. Codes are validated live in the top-up flow before a customer pays.
  • [Feature] Know-Your-Customer Verification - Optionally require identity verification before a customer can add funds or cross a spending threshold, so operators in regulated markets can meet their compliance obligations.
  • [Feature] Revenue Dashboard and Reports - A dedicated billing dashboard summarizes top-ups, consumption, taxes, and refunds. Detailed revenue and transaction reports can be exported for accounting, and a daily income digest email lands in the operator's inbox every morning.
  • [Feature] Billing Module SDK - A documented module framework lets developers add new payment gateways and ship custom billing front-end panels without modifying the core application.
  • [Feature] Self-Service Plan Changes - Customers can upgrade or downgrade an instance to a different plan directly from the instance settings. The instance is resized and the hourly rate is re-rated automatically, with the change prorated to the hour so the customer is only ever billed for what they used at each rate.
  • [Feature] AI Assistant, Now Manages Your Whole Stack - The built-in assistant moves well beyond answering questions. Customers can now ask it, in plain language, to create and manage instances, private networks and subnets, NAT gateways, load balancers, VPN gateways, managed databases, and Kubernetes clusters and node pools, with guided step-by-step flows that confirm the details before anything is built.
  • [Feature] AI Assistant for Operators - Administrators get a parallel set of assistant capabilities covering the platform itself: creating and modifying hypervisors and hypervisor groups, storage and backup configuration, service plans and plan groups, DNS providers, security groups and IP sets, object-storage servers and plans, currencies and credits, and users, roles, email templates, images, and access keys.
  • [Feature] Comprehensive Admin REST API - A large, consistent administrative API now covers compute, networking, storage, backups, billing, Kubernetes, databases, DNS, users and roles, and platform settings. Build custom automation, dashboards, and integrations against the same operations the admin panel uses.
  • [Improvement] Usage Report Export - The Cloud Service usage report can now be exported for offline analysis and accounting. This fulfills the export option that was previously marked as coming soon.
  • [Improvement] Streamlined Admin Navigation - The admin sidebar has been reorganized so the most-used sections sit at the top, and section headings are now searchable, making it faster to jump straight to a feature on a busy install.
  • [Improvement] Consolidated Billing Settings - All billing thresholds, currency display preferences, and suspension rules now live in a single settings surface, removing a long-standing source of confusion where related settings were split across two places.
  • [Improvement] AI Assistant Conversation Memory - Multi-turn conversations now correctly carry context from one message to the next, so follow-up questions and step-by-step deploy flows continue smoothly instead of starting over each turn.
  • [Improvement] AI Assistant Pricing Clarity - Credit balances and resource pricing shown by the assistant are now always denominated in the customer's selected currency, so estimates match what the customer sees everywhere else.