Grid Hosting maintenance (Resolvido)
  • Prioridade - Alta
  • Afetando Sistema - Grid Hosting Email Infrastructure
  • We will be performing a maintenance of our email infrastructure for Grid Hosting plans. This maintenance involves a few components, including upgrading the mail software to a new major version, along with an IP change of the underlying infrastructure.

    As a result of this, we will be temporarily take the email services offline. All inbound email will be queued by our inbound spamfilter servers, and be delivered once the system is back online. Outbound email will fail during the duration of the downtime, since the authentication service will be offline.

    We plan to start the work at around 10pm Amsterdam timezone on Wednesday. While we expect the downtime to only be roughly 5 minutes, we do allow a slightly longer window of downtime.

    Depending on some additional testing we're performing, we may have to postpone it by 24 or 48 hours. But the maintenance involves security fixes, so we'd like to get them done as soon as possible.

    Update 21:50: We'll begin at 22:00.

    Update 22:03: We're starting with taking the authentication and mail system offline.

    Update 22:06: The traffic has been redirected

    Update 22:11: The system has been updated to receive traffic again

    Update 22.13: The SSL certificates has been reissued

    Update 22:19: A quota permission issue has been resolved, that resulted in mail-accounts not updating quota

    Update 22:29: Final checks are being performed as we speak, and a change for hosting-panel.net are being rolled out to support the changes we've done

    Update 22:42: A final issue were spotted, which resulted in an invalid certificate being expired, which prevented most mail-clients from logging in without showing a certificate warning.

    Update 22:58: Based on logs, everything seems to be functioning as it should.

    SMTP services were down for 5 minutes and 19 seconds
    IMAP services were down for 7 minutes and 23 seconds
    Dovecot authentication were down for 7 minutes and 23 seconds (Same as IMAP services)
    Postfix authentication were down for 35 minutes and 6 seconds due to the reissued certificates were not correctly loaded into postfix due to wrong map type

    Maintenance concluded.

  • Data - 02/09/2026 22:00 - 02/09/2026 22:58
  • Ultima atualização - 02/09/2026 23:03
Emergency patching of shared infrastructure (Resolvido)
  • Prioridade - Crítico
  • Afetando Sistema - nlsh04,nlsh05,nlsh06,nlsh07,nlcp01,nlcp02,nlcp03,nash02
  • We will perform a reboot of servers (nlsh04 -> nlsh07, nash02, nlcp01 -> nlcp03) due to two security vulnerabilities: Bad Epoll (CVE-2026-46242) and GhostLock (CVE-2026-43499)

    Expected downtime is usually a couple of minutes to 15 minutes.

    servers nlcp05 -> nlcp08 will be at a later date.

    Updates:

    nash02: Updated, 6 minutes downtime. This will however receive another reboot later in the night as well, due to yet another vulnerability being discovered and patched later.

    We expect all other servers to get both patches as we reboot, since it will be late enough in the day.

    Update 15.52:
    nlcp05 -> nlcp08 are not vulnerable to Bad Epoll, and GhostLock is patched via KernelCare.

    Update 20.54:
    We will start rebooting the systems at 22.00.

    Update 21.55:
    We are starting with nlsh07 at 22.00

    Update 22.00: nlsh07 rebooting

    Update 22.04: nlsh07 successfully online again.

    Update 22.07: Proceeding with nlsh06

    Update 22.16: nlsh06 successfully online again.

    Update 22.17: Proceeding with nlsh05

    Update 22.25: nlsh05 successfully online again.

    Update 22.26: Proceeding with nlsh04

    Update 22.33: nlsh04 successfully online again.

    Update 22.36: Proceeding with nlcp03

    Update 22.42: nlcp03 successfully online again.

    Update 22.46: Proceeding with nlcp02

    Update 22.52: nlcp02 successfully online again.

    Update 23.05: Proceeding with nlcp01

    Update 23.11: nlcp01 successfully online again.

    This means we've completed the emergency maintenance of the servers in this batch. nash02 will be rebooted during night time in US.

    Update 06.27: nash02 is postponed until Wednesday morning.

    Update 06.34: We'll proceed with nash02

  • Data - 10/07/2026 12:09 - 11/07/2026 02:00
  • Ultima atualização - 16/07/2026 14:11
cPanel/WHM services closed (Resolvido)
  • Prioridade - Crítico
  • Afetando Sistema - cPanel/WHM services
  • Due to newly disclosed vulnerabilities in cPanel/WHM (CVE-2026-29201, CVE-2026-29202, and CVE-2026-29203), we are temporarily restricting access to all cPanel-related services. This measure is being taken proactively to protect customer environments until official patches have been released and verified.

    What remains operational
    All core services continue to function as normal:

    • Websites
    • Email (IMAP/SMTP)
    • FTP
    • SSH
    • All Grid Hosting plans

    Webmail users
    Customers who typically access email via webmail are advised to temporarily configure their email accounts in a local email client (e.g., Outlook, Apple Mail, Thunderbird).

    Email configuration details
    You can use the following settings to access your email:

    • Incoming mail server (IMAP): mail.yourdomain.com
    • Outgoing mail server (SMTP): mail.yourdomain.com

    Example: For myawesomewebsite.com, use mail.myawesomewebsite.com

    • IMAP port: 993 (SSL/TLS)
    • SMTP port: 465 (SSL/TLS) or 587 (STARTTLS)
    • Username: Your full email address
    • Password: The password associated with your email account

    These credentials are identical to those used for webmail login.

    Important notice
    We do not make exceptions during security lockdowns. Access to cPanel services will be restored as soon as patches are available and systems are secured.


    Updates

    • 10:17 AM: CVE identifiers added
    • 1:43 PM: Email configuration details added
    • 6.14 PM: cPanel released the software update a couple of minutes ago. We're applying the updates, which will take a bit of time. After that, we will re-enable cPanel and webmail interfaces.
    • 9.38 PM: We've re-enabled the webmail and cPanel interface. The case is hereby closed.

  • Data - 08/05/2026 08:52 - 08/05/2026 21:38
  • Ultima atualização - 08/05/2026 21:39
Access to cPanel blocked (Resolvido)
  • Prioridade - Crítico
  • Afetando Sistema - cPanel servers
  • Due to an active authentication bypass exploit found in cPanel. We've blocked access to cPanel for the time being until cPanel has rolled out a fix. The block means the control panel will not be accessible for the time being. Websites remain available.

    Update 10.38pm: The estimate is that cPanel will roll out the update during the night. We will enable access to cPanel (and webmail) again when we confirm it's indeed rolled out on servers. We expect this to be confirmed in the morning.

    We do recommend people who use webmail actively to set it up in a mail-client for the time being. For cPanel this usually boils down to mail. as IMAP and SMTP host. 993 for IMAP, 465 for SSL or 587 for StartTLS for SMTP. Your username is your email and password being the password for the email.

    https://support.cpanel.net/hc/en-us/articles/40073787579671-Critical-Vulnerability-with-cPanel-WHM-Login-Authentication 

    Update 07.00am: We've confirmed the update has been rolled out to all our cPanel servers. We've opened up for the services again.

  • Data - 28/04/2026 21:02 - 29/04/2026 07:00
  • Ultima atualização - 29/04/2026 07:02
nlsh04 + nlsh05 instability (Resolvido)
  • Prioridade - Crítico
  • Afetando Sistema - nlsh04+nlsh05
  • We experienced a few minor outages on servers nlsh04 and nlsh05.

    The outages were the result of rapid restarts of the webserver (litespeed) when updating mod_security rulesets from Imunify360. Due to a bug in the restart logic of Imunify360, this triggered rapid forceful restarts. This in turn results in the webserver itself stopping and starting with 1 minute delay.

    A fix has been implemented and deployed to all servers (nlsh04, nlsh05, nlsh06, nash02) to prevent this from happening, we will however have to monitor the situation for the next week or so to ensure all cases are handled.

  • Data - 08/04/2026 12:46 - 08/04/2026 15:47
  • Ultima atualização - 08/04/2026 15:50
nlsh04 NIC replacement (Resolvido)
  • Prioridade - Crítico
  • Afetando Servidor - nlsh04.h4r-infra.net
  • We'll have to replace the NIC in nlsh04 due to instability issues. This will be done likely Friday or Saturday.

    Update 28/Mar 9.21pm: We will replace the NIC tomorrow evening (sunday) around 10pm. This will result in a bit of downtime, since we have to power off the system, remove it from the rack, to switch out the hardware, rack it and turn it on again. After this we have to adjust network configs. We'll do our best to keep the downtime as short as possible.

    Update 29/Mar 8.18pm: We'll start the hardware replacement slightly after 10pm. We'll initially prepare new configs and the parts so we keep the downtime as low as possible.

    Update 29/Mar 10.23pm: We'll shut down the server in roughly 10 minutes.

    Update 29/Mar 22.50pm: The system is back online again, total downtime for IPv4 was 12 minutes and 7 seconds. IPv6 being 14 minutes and 9 seconds. We'll monitor the situation for a bit, before we leave the datacenter.

  • Data - 26/03/2026 07:02 - 08/04/2026 15:47
  • Ultima atualização - 08/04/2026 15:47
nlsh04 unavailability (Resolvido)
  • Prioridade - Crítico
  • Afetando Servidor - nlsh04.h4r-infra.net
  • We're experiencing a failure with server nlsh04.

    The downtime started at 06:16 amsterdam time.

    The cause is due to a failed network card, which requires replacement. We have a spare network card on-site for the replacement. A case has been created with the datacenter to perform the replacement.

    Update 06:57, we managed to get the system back online.

    We will schedule a hardware replacement in the coming days early mornings or late in the evening.

  • Data - 26/03/2026 06:16 - 08/04/2026 15:47
  • Ultima atualização - 26/03/2026 07:02
.dk domain transfers (Resolvido)
  • Prioridade - Baixa
  • Afetando Sistema - Registrar Services
  • All inbound .dk transfers are currently halted, while we're making some system changes. We expect this to be functional again somewhere between November 10 and november 20.

  • Data - 05/11/2025 08:00 - 26/03/2026 06:43
  • Ultima atualização - 05/11/2025 08:21
Server maintenance (Resolvido)
  • Prioridade - Média
  • Afetando Sistema - nlcp01,nlcp02,nlcp03,nlcp05,nlcp06,nlcp07,nlcp08,nlsh04
  • Over the coming week from Friday the 13th at 10pm Amsterdam timezone until Friday evening the 20th at 11:59pm we'll perform maintenance on multiple servers.

    To increase the redundancy we'll configure all servers in a bonded network configuration, what this effectively means is we double the network capacity of each individual server, but also ensure that even if one of our switches were to have issues, the other one will simply handle the traffic automatically with failover in place.

    We expect the downtime to be minimal for each server, however an outage of up to 15 minutes could be possible. In most cases the downtime will be less than 1 minute.

    13/Jun/2025:

    Update 9:58pm: We'll start with nlsh04 shortly

    Update 10:08pm: nlsh04 completed. IPv4 had 26 seconds of downtime, and IPv6 had 5 minutes and 12 seconds of downtime.

    Update 10:24pm: We'll proceed with nlcp01 shortly

    Update 11:16pm: An attempt was made to change to bonding on nlcp01, however it resulted in no connectivity after the bonding got re-established despite the bonding itself came up as expected. Multiple things were tried to restore connectivity, until we then decided to roll back the configuration to the original config. This also means we'll halt the changes for all remaining systems today.

    We'll try to replicate the issue in our test environment to see if we can reproduce the issue specifically related to this.

    14/Jun/2025:

    Update 01:02am: We managed to replicate the issue in our testing environment, specifically to specific cPanel configuration. A fix has been identified and implemented across all servers, and we will continue to perform the network change tomorrow evening.

    Update 11:17pm: We'll start with nlcp02 shortly.

    Update 11:32pm: nlcp02 completed. IPv4 had 26 seconds of downtime. We'll prepare nlcp01.

    Update 11:38pm: nlcp01 completed. IPv4 had 25 seconds of downtime. We'll prepare nlcp03.

    Update 11:57pm: nlcp03 completed. IPv4 had 33 seconds of downtime. We'll prepare nlcp05.

    15/Jun/2025:

    Update 00:18am: nlcp05 completed. IPv4 had 62 seconds of downtime. We'll prepare nlcp06.

    Update 01:08am: nlcp06 completed. IPv4 had 43 seconds of downtime. The remaining two systems will be completed in the evening (sunday) after 10pm.

    Update 9:55pm: We'll start with nlcp07 shortly.

    Update 10:10pm: nlcp07 completed. IPv4/IPv6 had 21 seconds of downtime. We'll proceed with nlcp08 shortly.

    Update 22:15pm: nlcp08 completed. IPv4 had 25 seconds of downtime. This marks the completion, and all servers now have bonded NICs.

  • Data - 13/06/2025 22:00 - 15/06/2025 22:15
  • Ultima atualização - 15/06/2025 22:16
nlcp03 emergency reboot (Resolvido)
  • Prioridade - Crítico
  • Afetando Servidor - NLCP03
  • We have to perform an emergency reboot of server nlcp03 due to a kernel issue.

    This reboot will be done around 10.15pm Amsterdam timezone. We expect around 5 minutes of downtime for the reboot, it may be a bit longer.

    The server was rebooted at 10:18pm and came back online again at 10.26pm.

  • Data - 09/05/2025 22:00 - 09/05/2025 22:26
  • Ultima atualização - 13/06/2025 10:13
nlcp03 downtime (Resolvido)
  • Prioridade - Crítico
  • Afetando Servidor - NLCP03
  • nlcp03 experienced some downtime during the evening, lasting on and off for a total of roughly 1 hour.

    Below you'll find a detailed event log:

    March 05:

    8.34pm: Our monitoring triggered an alert that nlcp03 were unavailable. Upon investigation we saw that the network driver for the Mellanox network card in the server had locked up.

    8.36pm: The once again became reachable and started serving traffic.

    9.55pm: We started seeing stability issues again, with random timeouts for some connections, while other succeeded. We continued to investigate possible causes and solutions.

    9.58pm: This was further confirmed by a customer ticket stating slow loading times as well.

    10.05pm: Due to the stability getting worse and worse, we restarted the system, due to the initial driver error, it can put the system into a state where drivers may not fully recover, and reboots may quite often solve this due to unloading and loading the network driver from scratch again.

    10.10pm: The server came online again.

    11.49pm: The system once again started causing stability issues, to resolve this, we implemented a possible fix by setting `iommu=pt` as one of the boot options, since this has been known to fix similar NIC errors on AMD EPYC 7002 series servers (Which is what the particular server runs).

    11.51pm: The system came online, we began our verification.

    March 06:

    12.05am: The iommu=pt fix didn't get implemented correctly, due to tuned profiles overriding the options, we changed the implementation and rebooted the system again to apply the iommu=pt fix.

    12.10am: The system came online, and we began verifying the fix.

    12:31am: The iommu=pt fix didn't resolve the issue, making us believe that the cause is a failing network card, we prepared a replacement server, and started our route towards the datacenter.

    12:52am: Another reboot was made to bring back the system temporarily after losing the network.

    01.15am: We arrived at the datacenter, passed through security, ensured the switch configurations were correct for the replacement hardware.

    01:36am: nlcp03 powered down to move the drives to another physical server

    01.41am: The server came back online in the new chassis, and we're since monitoring the situation.

    The old system will be receiving some temporary drives, so we can perform some additional tests.

  • Data - 05/03/2025 20:35 - 06/03/2025 02:00
  • Ultima atualização - 07/03/2025 23:50
Datacenter migration (Resolvido)
  • Prioridade - Alta
  • Afetando Outro - nlcp01,nlcp02,nlcp03,nlcp05,nlcp06,nlcp07,nlcp08,nlsh04,proxmox05,proxmox06,proxmox07
  • Between 10pm the 18th, and 2am the 19th of December, a datacenter migration took place.

    Most systems were online within 2 hours (so around midnight).

    nlcp06 had some issues, which took longer to resolve due to a broken PSU, the PSU was replaced with an on-site spare PSU.

    nlsh04 had issues due to the per-customer CPU, memory and disk limits were not being applied correctly. After we brought some internal systems up, this started to work again. We've taken steps to ensure this will not happen in the future.

    During January and February, we'll slowly increase the redundancy on the network config on each server, this requires reconfiguring the servers network interfaces, as well as the switches, to apply the redundancy.

    This will be done during the night after midnight, and it will in some cases cause up to 5-10 minutes of downtime, depending on the system configuration, in many cases however, it will be less than a minute.

  • Data - 18/12/2024 22:00 - 19/12/2024 02:00
  • Ultima atualização - 19/12/2024 03:45
hosting-panel.net redirector and URL Scheduler unavailable (Resolvido)
  • Prioridade - Média
  • Afetando Sistema - Redirector and URL Scheduler
  • As of 2am Amsterdam timezone, our Redirector and URL Scheduler for hosting-panel.net are unavailable. These two systems are located with the same upstream provider. A solution is being worked on as we speak.

    3:10am: The Redirector is back online. The URL Scheduler remains unavailable at this time.

    9.55am: The system is back online.

  • Data - 02/09/2024 02:00 - 02/09/2024 09:55
  • Ultima atualização - 02/09/2024 16:18
.dk domain registrations delayed (Resolvido)
  • Prioridade - Alta
  • Afetando Sistema - .dk domain registrations
  • We're aware of delays in .dk processing / registration

    We're investigating the issue with our suppliers, but it boils down to a change in how data is required in the EPP service for .dk domains. We're working on resolving it as soon as possible.

    Update 18/04/2024 1:30pm Amsterdam timezone: The problem has been resolved

  • Data - 17/04/2024 14:00 - 18/04/2024 13:30
  • Ultima atualização - 18/04/2024 13:30
de-mail01 outage and replacement (Resolvido)
  • Prioridade - Crítico
  • Afetando Sistema - de-mail01
  • At 10.50pm we got a report of slow email sending on de-mail01

    At 11.24pm after investigating the issue we decided to reboot the system due to indicators that the systemd process (the main management process) didn't function as it should causing various issues with managing systems.

    As a result of this we performed a manual backup of the system, including email and configuration.

    At 11.42pm the system was restarted and came back online as normal after a few minutes.

    We decided at the same time to do a disk integrity check due to the fact we had to force reboot the system. The check indicated multiple errors on both drives in the storage array.

    Due to both disks being affected, the #1 priority were to get a new system online, and configure it to take over the email accounts located on the system.

    At 01.05am We started the testing of the configured system

    At 01.34am We stopped dovecot and postfix on the old system, performing the last migration of files.

    At 01.46am We switched the DNS, updated inbound mail-routing and updating the records in hosting-panel.net to reflect the new system.

    At 01.50am We performed an update to fix some permissions to restore correct permissions to all accounts.

    At 02.15am We completed our checks.

    Backups has been re-enabled for the new system, the system is configured for monitoring.

    The only last thing that will be done during tomorrow is restoring the full text search engine to speed up full text searching in dovecot.

  • Data - 21/01/2024 22:50 - 22/01/2024 02:15
  • Ultima atualização - 22/01/2024 02:28
NLSH04 unavailability (Resolvido)
  • Prioridade - Crítico
  • Afetando Servidor - nlsh04.h4r-infra.net
  • Between 1.55AM and 3.31AM UTC we experienced a total of 15 minutes of downtime on server NLSH04.

    This downtime was the cause of an unplanned urgent reboot due to a broken kernel module.

    While provisioning new users on the system, we realized that some overall account limits were not applied correctly, after investigating this, we found out that a kernel module managing these limits were not loaded correctly. Due to the nature of the module it cannot be loaded again without rebooting the system.

    The initial reboot we performed resulted in the module being loaded correctly, the downtime for this reboot was 5 minutes per our monitoring.

    We however discovered that certain CloudLinux LVE features were no longer available due to a failure that caused certain software features to not work as expected. This prompted a second reboot, however, because the configuration was updated as a part of the software update, it resulted in critical boot parameters to no longer be present (parameters that effectively disables cgroups v2 which kmod-lve from CloudLinux is incompatible with). With the module not loading, it causes the system to not being able to make users enter a virtual secure environment, effectively rendering the service unavailable. Fixing these parameters and performing one final reboot, caused the system to come back up as expected.

    The downtime for the 2nd outage lasted for 10 minutes.

    We'll be taking additional steps to implement additional metrics into our systems to catch this faster, which hopefully should prevent an issue similar to this to occur in the future.

    We're sorry for the inconvenience caused by the 15 minutes downtime. 

  • Data - 18/11/2023 02:55 - 18/11/2023 04:31
  • Ultima atualização - 18/11/2023 04:09
Network outage (Resolvido)
  • Prioridade - Crítico
  • Afetando Sistema - NLCP05, NLCP07
  • We're currently experiencing network connectivity issues on the servers NLCP07 and NLCP05 due to loss of connectivity, we're investigating the cause of this.

    Update 4.55pm: The issue has been resolved after 41 minutes of downtime.

    We're still looking into the definitive root cause of the problem, why it triggered in the first place, however, our current findings:

    servers nlcp05, nlcp07 as well as a 3rd server used for some internal services lost their IPv4 connectivity at 4.14pm.

    Upon investigation we saw that the switch side had lost the "arp" entries for IPv4, meaning the switch would stop knowing how to route packets. We however, saw that IPv6 worked without issue.

    Ultimately we decided to disable a protection on the switch in regards to arp, which restored connectivity immediately for all systems.

    We're still investigating the actual root cause, but for now, the systems should be functioning.

  • Data - 06/08/2023 16:22 - 06/08/2023 16:55
  • Ultima atualização - 06/08/2023 17:15
Software + Hardware upgrade of all cPanel hosting (Resolvido)
  • Prioridade - Média
  • Afetando Outro - NLCP01 to NLCP08
  • For continued security and stability, we'll perform necessary software upgrades that require a system reboot. However, to further increase performance and capacity, we'll also upgrade our systems with additional memory and storage space.

    This means we will take each server offline for up to 30 minutes (the expected time being much less in reality). We'll perform the upgrade between 3.30am and 7.30am. We'll do a single server at a time, in case we do not manage to upgrade all 8 servers in one go, we'll schedule another timeframe for the remaining servers shortly after.

    The software maintenance is required to ensure the continued security of our platform. We strive to keep the downtime as low as possible.

    Update Aug4 1:23am: We'll start the maintenance at 3.30am

    Update Aug4 3:57am: We'll start with nlcp08

    Update Aug4 4.23am: nlcp08 up with 10 minutes of downtime. Proceeding with nlcp07 in a couple of minutes.

    Update Aug4 4.36am: nlcp07 up with 9 minutes of downtime. Proceeding with nlcp06 in a couple of minutes.

    Update Aug4 4.52am: nlcp06 up with 11 minutes of downtime. Proceeding with nlcp05 in a couple of minutes.

    Update Aug4 5.05am: nlcp05 up with 11 minutes of downtime. We'll take a short break and continue with the last 4 shortly.

    Update Aug4 5.23am: Proceeding with nlcp04

    Update Aug4 5.36am: nlcp04 up with 12 minutes of downtime. Proceeding with nlcp03 in a couple of minutes.

    Update Aug4 5.54am: nlcp03 up with 8 minutes of downtime. Proceeding with nlcp02 in a couple of minutes.

    Update Aug4 6.11am: nlcp02 up with 8 minutes of downtime. Proceeding with nlcp01 in a couple of minutes.

    Update Aug4 6.25am: nlcp01 up with 10 minutes of downtime. Maintenance complete.

  • Data - 03/08/2023 03:30 - 03/08/2023 07:30
  • Ultima atualização - 04/08/2023 06:26
Reduced backup rotation (Resolvido)
  • Prioridade - Média
  • Afetando Sistema - backup
  • We're running with slightly reduced backup rotation while moving some systems around. Due to a fault in the backup system, we've had to shift traffic to a different system to then redo the main backup system. This means there's a limited number of days available (still above 7 days of data).

  • Data - 17/06/2023 18:15 - 17/07/2023 00:00
  • Ultima atualização - 04/08/2023 01:26
System crash (Resolvido)
  • Prioridade - Crítico
  • Afetando Servidor - NLCP01
  • Between 03:45 and 03:55, we experienced 10 minutes of downtime on nlcp01.

    The cause of the crash based on our investigation seems to be a kernel lockup, which resulted in the system freezing.

    To resolve the issue a hardware reset were performed and the system came online as expected.

    We have a planned maintenance for later this month to upgrade the kernel among other software, which should likely resolve the these issues.

    We're sorry for the inconvenience caused by this.

  • Data - 09/07/2023 03:45 - 09/07/2023 03:55
  • Ultima atualização - 09/07/2023 04:03
Jetbackup 5 restoration (Resolvido)
  • Prioridade - Baixa
  • Afetando Outro - Backup - jetbackup 5 - nlcp05/nlcp06
  • Customers on nlcp05 and nlcp06 might experience issues after restoring a backup via JetBackup 5.

    The issues has been reported to Jetapps (The developers of JetBackup), and we're waiting for a release to fix the issues.

    404 page after restore of subdomains or addon domains

    JetBackup restores incorrect folder permissions for the "document root" of the domain.

    The permissions are supposed to be "0755" but are restored as "0750" - this can be corrected after restoration in File Manager or FTP by selecting the folder for the domain and change the permissions.

    User should have read/write/execute permissions.

    Group and World/Everyone should have read/execute permissions.

    You can also create a ticket for us to correct it.

    Database restoration requires user restore

    Jetbackup when restoring a database, does not restore user permissions to the database.

    You either have to restore the DB User during restore as well for the corresponding database (this ensures that the "grants" are restored as they should be).

    Alternatively, you can go to "MySQL® Databases" in cPanel, go to the "Add User to Database" section, select both the user and the database, click "Add" and assign all privileges to the database. This will restore the database permissions.

  • Data - 26/05/2021 10:44
  • Ultima atualização - 02/07/2023 17:00
Backup Server unavailable (Resolvido)
  • Prioridade - Alta
  • Due to a failure of a top of rack switch, our backup system is currently unavailable for the time being, we're working on getting the connectivity restored as soon as possible.

    When we can access the system again, it will be configured to have redundant uplinks.

    Update 20:29: The system is available again

  • Data - 01/02/2023 17:13 - 01/02/2023 20:29
  • Ultima atualização - 02/02/2023 01:16
Replacement of failing RAM (Resolvido)
  • Prioridade - Alta
  • Afetando Sistema - de-mail01
  • We have a DIMM in our de-mail01 system that is spitting out ECC correctable errors, since this usually will turn into uncorrectable errors over time, we're proactively replacing the memory of the server.

    Since memory can't be changed during runtime, we'll have to power off the system at 11pm this evening to make the datacenter change the memory (this should be a relatively short process), we can however expect anywhere from 10 to 30 minutes downtime as a result before all services return to normal.

    Our inbound mail-system will hold onto emails in the meantime for known email accounts on the given system.

    Accessing email accounts located on the server (hosting-panel.net related accounts) will be unavailable in the time being.

    Update 10.47pm: We'll shut down the system in a couple of minutes.

    Update 11.33pm: The system is back online

  • Data - 13/01/2023 23:00 - 13/01/2023 23:59
  • Ultima atualização - 13/01/2023 23:37
backup migration to JetBackup 5 (Resolvido)
  • Prioridade - Baixa
  • Afetando Sistema - nlcp01-nlcp04
  • Over the coming weeks we're slowly migrating the servers nlcp01, nlcp02, nlcp03 and nlcp04 from JetBackup 4 to JetBackup 5. Since these two versions are incompatible on a storage level, we're therefore keeping both systems "live" for 28 days per server.

    We've started the migration of nlcp01, nlcp02 and nlcp03 - you'll find both JetBackup 4 and JetBackup 5 in cPanel. Over time more and more recovery points will be available in JetBackup 5, and backups in JetBackup 4 will be rotated out.

  • Data - 09/10/2021 17:43 - 09/11/2021 17:43
  • Ultima atualização - 18/11/2022 01:38
Outage of nlcp02 (Resolvido)
  • Prioridade - Crítico
  • Afetando Servidor - NLCP02
  • Between 8.51am and 9.48am Amsterdam time today, we experienced a total of 12 minutes of downtime of nlcp02, some sites may have been up to 20 minutes total.

    This event was caused by two issues:

    • MySQL reported down which required us to forcefully kill the database
    • Subsequently the webserver decided to lock up, resulting in 503 Service Unavailable errors being returned after MySQL itself had recovered

    We were unable to diagnostic the actual issue of MySQL since we could not query the system for any information, trying to gracefully stop the system did not do anything, where we eventually had to kill it forcefully. It however, resulted in two databases ending up with 1 table each that were not in a healthy state.

    After MySQL was brought back online, we continued to experience roughly 5 minutes of issues with the webserver returning "503 service unavailable" errors. The cause for this was a request backlog that had to be processed for sites, so some sites would recover immediately where others would have been down for slightly longer, depending on how quick the queue got cleared.

    Two databases needed to get their tables repaired, which were successful, and without any corruption or loss for the given databases.

  • Data - 17/11/2022 08:51 - 17/11/2022 09:48
  • Ultima atualização - 17/11/2022 11:19
host reboot (Resolvido)
  • Prioridade - Crítico
  • Afetando Sistema - NLSH01
  • We're investigating an issue with the networking on nlsh01, which might in worst case require a reboot of the system.

    We're sorry for the inconvenience 

  • Data - 22/03/2022 09:22
  • Ultima atualização - 20/04/2022 21:47
Halting incoming migrations (Resolvido)
  • Prioridade - Baixa
  • Afetando Outro - Migrations
  • Between November 15th at 00:00 until December 17th 00:00, we won't perform any incoming migrations, neither the free ones or paid.

    All migrations that will involve us, will be planned either before or after this period.

  • Data - 15/11/2021 00:00 - 17/12/2021 00:00
  • Ultima atualização - 11/02/2022 13:18
Network maintenance - WorldStream (Resolvido)
  • Prioridade - Média
  • Afetando Sistema - nlcp01 - nlcp06
  • Between 6 am and 8 am on August 25th, WorldStream, the data center we use for all shared hosting servers, will perform network maintenance which requires updating the core routers and one distribution router to a new software release.

    During this timeframe, there may be a short moment (a matter of seconds) when traffic switches from one router to another. This can occur multiple times during maintenance.

    The network is fully redundant, and our systems are connected to two different routers, so impact remains minimal. It does put a higher risk on the uptime for the duration of the maintenance due to the lack of redundancy for some time.

    The maintenance page at WorldStream will be kept up to date as well: https://noc.worldstream.nl/ 

  • Data - 25/08/2021 06:00 - 25/08/2021 08:00
  • Ultima atualização - 14/09/2021 10:31
Network drop (Resolvido)
  • Prioridade - Alta
  • Afetando Sistema - nlcp01 - nlcp06
  • Between 00:01 and 00:03 we experienced a network drop affecting all servers in the WorldStream datacenter, due to a connectivity issue between WorldStream and Nikhef (A core network location in Netherlands) - the drop lasted for less than a minute.

  • Data - 24/08/2021 00:01 - 24/08/2021 00:03
  • Ultima atualização - 24/08/2021 11:23
Network maintenance - WorldStream (Resolvido)
  • Prioridade - Média
  • Afetando Sistema - nlcp01 - nlcp06
  • Between 6 am and 8 am on August 4th, WorldStream, the data center we use for all shared hosting servers, will perform network maintenance which requires updating the core routers and one distribution router to a new software release.

    During this timeframe, there may be a short moment (a matter of seconds) when traffic switches from one router to another. This can occur multiple times during maintenance.

    The network is fully redundant, and our systems are connected to two different routers, so impact remains minimal. It does put a higher risk on the uptime for the duration of the maintenance due to the lack of redundancy for some time.

    The maintenance page at WorldStream will be kept up to date as well: https://noc.worldstream.nl/ 

  • Data - 04/08/2021 06:00 - 04/08/2021 08:00
  • Ultima atualização - 18/08/2021 14:46
Kernel Upgrade - reboot required (Resolvido)
  • Prioridade - Alta
  • Afetando Sistema - All servers
  • On July 29th at 10.30 pm, we will perform a kernel update of all servers. This update is considered urgent due to the vulnerabilities CVE-2021-22555 and CVE-2021-33909.

    We do expect up to 10-15 minutes of downtime per server.

    The kernel is scheduled for release on July 27, so we're making room for a slight delay in the update being made available.

    Update 29/07: The kernel update for servers nlcp01, nlcp02, nlcp03, and nlcp04 have been postponed until August 4th due to a discovered bug in the el7h kernel from CloudLinux. They're releasing a fix for this today, however, they only expect the full rollout of the kernel on August 3rd.

    Servers nlcp05 and nlcp06 will still be updated today, since these rely on the el8 kernel which does not have this bug.

    Update 29/07 10.29pm: We're starting with nlcp06

    Update 29/07 10.36pm: We've completed nlcp06, starting nlcp05 shortly

    Update 29/07 10.46pm: We've completed nlcp05.

    Downtime for each server being about 5 minutes. nlcp01 to nlcp04 will be rebooted next week.

    Update 04/08 10.23pm: We're starting with nlcp04 shortly.

    Update 04/08 10.37pm: nlcp04 done, we're proceeding with nlcp03 shortly.

    Update 04/08 10.47pm: nlcp03 done, we'll proceed with nlcp02 shortly.

    Update 04/08 10.53pm: nlcp02 done, we'll proceed with nlcp01 shortly.

    Update 04/08 10.59pm: nlcp01 done - all servers had a downtime of 3-4 minutes. Maintenance completed.

  • Data - 29/07/2021 22:30 - 05/08/2021 03:00
  • Ultima atualização - 04/08/2021 22:59
Network maintenance (Resolvido)
  • Prioridade - Média
  • Afetando Outro - DC1 - Dronten - internal + managed VMs
  • Between 4 AM and 6 AM on the 27th of July, there will be performed a network maintenance in the Dronten datacenter managing some internal systems as well as some managed customers.

    This is a part of emergency maintenance work to resolve stability issues with the network.

    A total downtime of up to 1 hour can be expected during this timeframe.

  • Data - 27/07/2021 04:00 - 27/07/2021 06:00
  • Ultima atualização - 29/07/2021 10:30
systems unavailable - crash (Resolvido)
  • Prioridade - Crítico
  • Afetando Sistema - nlcp01, nlcp02, nlcp05
  • At 10:15 we saw nlcp01 and nlcp02 becoming unavailable, and nlcp05 became unavailable at 10:19.

    The systems were rebooted and the last came online at 10:24.

    The cause of the crash is due to KernelCare live patching which triggered a so-called kernel panic.

    We've disabled automatic patching for the time being and have informed KernelCare about the issues.

    KernelCare has disabled the patch-set and estimates a fixed patch in the coming week.

  • Data - 22/07/2021 10:15 - 22/07/2021 10:24
  • Ultima atualização - 22/07/2021 13:29
backup schedule reduced (Resolvido)
  • Prioridade - Média
  • Afetando Sistema - NLCP05,NLCP06
  • We're temporarily decreasing the backup schedule from every 6 hours to once per day.

    Due to a bug in how backups are linked together to produce incremental backups, this currently doesn't function correctly.
    The result of this being a full backup is taken on every run, using up 700GB of disk space every 6 hours, which doesn't scale, we'd run out of disk space on our storage server fairly quickly.

    We are, however, going to configure a second backup job, to perform backups of the databases every 6 hours for the time being.

    As of 2.30pm, we've enabled full rotation again.

  • Data - 04/05/2021 15:33 - 05/05/2021 14:30
  • Ultima atualização - 05/05/2021 15:59
DNS unavailable (Resolvido)
  • Prioridade - Crítico
  • Afetando Sistema - DNS
  • Between 15:59 and 16:02 we experienced complete unavailability of our DNS infrastructure due to an attack on a particular domain.

    We've since implemented some additional measures to try to mitigate it in the future. We're still investigating why the attack happened in the first place.

    We're sorry about the inconvenience caused by the outage.

  • Data - 16/04/2021 15:59 - 16/04/2021 16:02
  • Ultima atualização - 16/04/2021 17:11
Reconfigure Backup Server (Resolvido)
  • Prioridade - Baixa
  • Afetando Sistema - Backup server
  • Due to a recent change in our backup configuration which requires a lot more inodes than previously, we're required to reformat the backup system to use another filesystem that allows for a larger number than what ext4 gives with the current 20TB of storage we have.

    As a result of that, as of today, we're performing backups to a separate location than usual (Two smaller physical servers), these servers are less powerful than the current system, meaning backups and restores may be slowed down slightly as a result of this.

    These two servers will be used as a temporary storage location for new snapshots.

    Normally we store 28 days of recovery points - as a part of this backup maintenance, we'll lower the number to 14 days, so we can get back to the original storage system again within a decent timeframe (14 days instead of 28 days).

    If you do wish to keep a few additional snapshots for longer than this period, then you'll have to download these snapshots via cPanel within 14 days.

    This maintenance is strictly required since we're nearing the limitations of the current inode count on the current filesystem.

  • Data - 12/09/2020 12:09 - 15/01/2021 15:08
  • Ultima atualização - 15/01/2021 15:08
Acronis unavailable on some systems (Resolvido)
  • Prioridade - Média
  • Afetando Sistema - Acronis Backup
  • We're investigating issues with Acronis recovery points being unavailable on some systems.

    We're able to do restorations of files (not databases) via the Acronis Console directly - we've upped the JetBackup storage time from 7 days to 28 days, and enabled 6-hour snapshots in JetBackup as well, in the meantime.

    While Acronis continues to perform backups, we're not able to restore them directly from cPanel, this also limits it to only files being able to be restored. Databases can be restored in a disaster recovery scenario, but it's not an easy task.

    Please use JetBackup for the time being for performing restorations, while we're recovering the functionality in cPanel.

    nlcp01 to nlcp10 and server9 are unaffected since they're using JetBackup by default as the main backup functionality.

    ETA for resolution is currently unknown.

    Update 7.18pm: After investigation together with the Acronis support and a bunch of debugging, the result so far is that some of the disksafes are corrupted after an attempted repair.

    As a result of this, new disksafes has been made and are backing up again, however the recovery points prior to today are lost. All servers where the safes has been deleted, JetBackup has been doing backups cleanly, so recovery is still possible, however at a smaller timeframe than usual.

    We're trying to get the disksafe for server16 to work properly, since in this particular case, we're only using Acronis backup - Jetbackup has been enabled on this server, however since there's only backups from today, recovery longer than that is currently not possible.

    Update 12:07am: server16 disksafe are rendered corrupted, thus recovery is not possible, we're still checking on server15 and server8.

    Update 8.24am: Disksafes server15 and server8 are only functional directly within Acronis console - so DB restores will only be available at the current Jetbackup rotation.

    Additionally, we've disabled new Acronis backups on server16, so only the old ones are available, Jetbackup is used as the primary backup source moving forward.

    We close the case, since backups are functional, despite the lack of some history on multiple servers - disaster recoveries are possible, and backups are available.

    nlcp01 to nlcp10 will only use Jetbackup, despite being a bit more resource-heavy, it's providing the most reliable restoration and storage capabilities.

  • Data - 24/08/2020 08:00 - 26/08/2020 08:30
  • Ultima atualização - 28/08/2020 10:56
server unavailable (Resolvido)
  • Prioridade - Crítico
  • Afetando Servidor - FRA16
  • Server16 is currently unavailable. Status can be followed at https://twitter.com/Hosting4Real/status/1285506453604900864 

    Update 13:06 - the system is back online

  • Data - 21/07/2020 11:25 - 21/07/2020 13:06
  • Ultima atualização - 21/07/2020 13:10
Email delivery issues (Resolvido)
  • Prioridade - Alta
  • Afetando Sistema - Spam Filter
  • Earlier today we received notifications from some customers that their customers informed them regarding DNS lookup failures.

    Initially, by the looks of the error message, it indicated a failure in the DNS resolution for the domains themselves, with the way our DNS servers are configured, it's unlikely this is caused by our DNS servers being unavailable (They're located in 4 different data centers, 4 different providers in 3 different countries). This first indicated a possible resolver issue at Microsoft due to the nature of it.

    Upon further investigation, we saw that the resolution error was not related to the domains themselves, but to the mail exchange DNS (The NDR report from Microsoft didn't actually indicate this).

    Testing our delivery we saw that the connections towards mx01.hosting4real.net and mx02.hosting4real.net would hang. Doing DNS lookups to resolve these two MX entries resulted in DNS timeouts.

    Further testing from multiple Microsoft Azure locations, we saw that the DNS provider (Zilore) we use for the domain hosting4real.net was not actually reachable from within Microsoft's network. The Zilore NOC team was informed about the findings, the result being only DNS queries routed to Zilore's South African datacenter would have issues (Microsoft for some reason were routed to South Africa).

    While Zilore worked out the issues with the DNS, we decided to implement a secondary MX domain into our spam filter solution, and start updating DNS entries for the customers affected by this.

    We've thus added mx01 and mx02.h4r.eu as DNS for our spam filter servers, the h4r.eu domain uses another set of DNS servers for resolution (our standard ones).

    Overall this should improve the availability even further due to the fact we a secondary DNS provider available for the spam filtering as well.

    We're sorry about the inconvenience caused by this, the far majority of the emails will be delivered by Microsoft since they retry, however, in case they've reached their maximum retry limit, this will cause a failure, and will require the sender to send the email again.

  • Data - 23/06/2020 12:40 - 23/06/2020 15:45
  • Ultima atualização - 23/06/2020 16:57
Servers unavailable (Resolvido)
  • Prioridade - Crítico
  • Afetando Outro - RBX Datacenters
  • We're experiencing issues in Roubaix datacenters for server7, server8, server13, and server15 - we're investigating.

    Update 5:34pm: The outage is caused by a network outage at the datacenters of Roubaix.

    Update 5:43pm: The network seems to have returned to normal. We're still waiting for a reason for the outage to be provided by the datacenter. We continue to monitor the recovery of traffic to the affected servers.

    Update 10.20pm: The network outage was caused by a router crash, the data center provider is investigating together with the network vendor to see what caused the issue. In the meantime parts of the router have been isolated and certain links (800gbps in total) have been reenabled to increase the capacity further towards Amsterdam and Frankfurt.

    When the full investigation has been completed and this is announced, we'll update the post.

    Update 11.10pm: As of 10.56pm the capacity has been increased to 2100gbps for the router.

    RFO:

    The root cause for the outage was caused due to a hardware failure of a daughter card (linecard) in the RBX-D1-A75 router, happening at the RAM parity level, this linecard originally raised alerts the 25th of March which the manufacturer confirmed at the 27th of March wouldn't be critical and should simply reboot the card during next maintenance.

    Monday the 30th of March the errors appeared again on the same card, leading to the corruption of the software and preventing isolation of the card. The failed card propagated the corruption within the router so it became unstable and caused it to crash.

    This means that 50% of the traffic passing through the Roubaix backbone would be affected since it would pass through this router pair.

    For improvement, the provider are working on creating isolated availability zones to reduce the impact even further if it should happen again.

    Additionally, regular redundancy tests will be performed at the backbone level.

  • Data - 30/03/2020 17:14 - 30/03/2020 17:43
  • Ultima atualização - 26/04/2020 20:12
Payment gateway (Resolvido)
  • Prioridade - Baixa
  • Afetando Sistema - Payment gateway
  • After the switch to Stripe, we saw that Nets the danish Card Issuer started to reject payments for our merchant ID - meanwhile Nets and Stripe are working on this, we've switched back to Braintree Payments to allow customers to pay with card again.

  • Data - 19/09/2019 09:31
  • Ultima atualização - 15/12/2019 21:45
Backup system reinstallation (Resolvido)
  • Prioridade - Baixa
  • Afetando Sistema - Backup system
  • Monday we'll reinstall our backup system due to an issue with the disk array that is unrecoverable, so we have to destroy the array and create it again (with new disks in it).

    In the meantime, we've enabled a secondary backup server to take over backups for the time being until the original system is back up and running.

    This also means we lose some backup history, we'll have backups for the beginning of October (Until the 7th or 8th of October), as well as starting from today the 25th.

    Restoring backups from the 8-9th until 24th will not be possible since the data will be gone.

  • Data - 28/10/2019 08:00
  • Ultima atualização - 15/12/2019 21:45
server7 / rbx7 unavailable (Resolvido)
  • Prioridade - Crítico
  • Afetando Servidor - RBX7
  • We're experiencing issues with server7.

    At 00.18 the system rebooted, the reason currently is unknown

    The system is copying a large file from /tmp to /usr/tmp.secure which is blocking services from starting.

    We're sorry for the inconvenience caused by this.

    Update: The server is online as of 00.57. We'll continue to investigate the root cause.

    Update: After investigating the issue, we believe it has been a correlation between a MySQL lock caused by our backup software and a system package being updated in the same process (which happens to also affect the backup software).

    We've went through our infrastructure to ensure that these two tasks doesn't run close to each other, as well as added additional logging to MySQL to see if it should happen in the future, we can see where the lock is caused.

  • Data - 03/10/2019 00:18 - 03/10/2019 00:57
  • Ultima atualização - 05/10/2019 17:27
Exchanging payment gateway (Resolvido)
  • Prioridade - Baixa
  • We're in the process of changing payment gateway from Braintree Payments to Stripe.

    We've tested the integration and we can confirm that it works as expected.

    As a part of our maintenance upgrading our billing system software to the new major version, we'll do the change of the gateway in the same maintenance window.

    We'll perform our maintenance Monday during the day. We expect the maintenance to last roughly 2 hours.

    Update 07:09: We'll start the maintenance

    Update 07.36: Maintenance has been completed and we have confirmed that our gateways can process payments as expected.

  • Data - 01/09/2019 20:20 - 02/09/2019 07:37
  • Ultima atualização - 02/09/2019 07:37
Server unavailable (Resolvido)
  • Prioridade - Crítico
  • Afetando Servidor - GRA10
  • We're currently experiencing issues with the availability of server10 - the datacenter is aware and working on a solution.

    Update 25/07/2019 3.19pm: Since it's affecting the whole rack, we're expecting it to be the top of rack switches (both public and IPMI) that has shut down due to the heatwave currently hitting France.

    Update 25/07/2019 3.31pm: The network has returned and the server is reachable again. We continue to monitor the situation. 

    Update 26/07/2019 07:14am: We've received the root cause of the problem as of 06.47am this morning.

    The issue that happened yesterday, was a result of a misconfiguration on some of the switch equipment within the data center where the maximum temperature configuration of the switch was simply set too low, this has been corrected for the switches that resulted in downtime, and over the coming days the DC provider will ensure the consistency of the settings across their 25 data centers and thousands of switches.

    OVH the provider we use for our main operation such as web hosting runs a quite unique setup when it comes to cooling data centers. Normal data centers are cooled by HVACs (Heating, ventilation, and air conditioning) , some will use a mix of HVAC and free cooling where you use a mix of the outside air to cool with if the air is cold enough. Some data centers will do water cooling in their racks using an indirect cooling method by having a loop in the rack that then chills the air to provide cold air for the servers.

    OVH does actual direct water cooling on their systems, meaning every server has its own water loop (connected to a bigger loop), additionally, they have to circuits for this operation, the remaining cooling in the data centers is done at a "per room" basis with 1 water circuit that cools the air in the rooms.

    Another loop will be added to the indirect air cooling system in every room, effectively doubling the capacity of the cooling system, and thus further lowering the temperature in the rooms.

    Data centers located in cities where high temperature and high humidity is a possibility, additional air cooling systems are installed to cope with the heat. Whether this is expanded to other cities is on a "site per site" basis.

    Early 2020, OVH will work on a new proof of concept to further improve the cooling capacities in their data centers.

  • Data - 25/07/2019 14:57 - 25/07/2019 15:31
  • Ultima atualização - 26/07/2019 07:27
false positive downtime alert (Resolvido)
  • Prioridade - Baixa
  • Afetando Sistema - server12,server14
  • At 4.52pm we received an alert about server14 being unavailable - however the server continued to receive traffic.

    After investigation we saw that the monitoring system got graylisted by our Imunify360 web application firewall, thus failing our test marking the server as down.

    At 4.59pm server12 alerted about downtime.

    At 5pm the servers got marked as "online" again, after we implemented a fix.

    The fix has been implemented on all servers to avoid these false positives in the future.

  • Data - 14/06/2019 16:52 - 14/06/2019 17:00
  • Ultima atualização - 14/06/2019 17:24
Reboot to fix kernel bug (Resolvido)
  • Prioridade - Média
  • Afetando Servidor - GRA14
  • We'll have to reboot server14 to fix a kernel bug - expected downtime will be roughly 5 minutes.

    update 10.01pm: We'll reboot the server in a minute.

    update 10.11pm: Server has been rebooted, total downtime being 5 minutes and 20 seconds.

  • Data - 09/06/2019 22:00 - 09/06/2019 22:11
  • Ultima atualização - 09/06/2019 22:11
reboot of all servers (Resolvido)
  • Prioridade - Alta
  • Afetando Outro - All servers
  • A recent vulnerability (Zombieload) in Intel CPUs, requires that we reboot all systems to install microcode updates to the CPU.

    We expect somewhere between 5 and 10 minutes of downtime per server.

    In rare cases there can be a boot problem, which will be resolved as quickly as possible, but the risk is there.

    This update comes at a short notice, but due to the severity of the vulnerability, it cannot wait.

    We're sorry about the inconvenience caused by this.  

    Update May 16: We'll be able to patch Zombieload without the need of reboots thanks to KernelCare. The patch is expected to arrive friday.

    Update May 17: We have to reboot server7 to server12 this evening due to the CPU version we're using in those servers. We'll do one server at a time, starting with server7 at 8pm.

    We're sorry for the inconvenience caused by this - however, the security of the systems is number 1 priority.

    We'll also have to migrate a few customers in the coming weeks to rebalance the CPU usage - those customers will be contacted.

    Update May 17 8.05pm: We're rebooting server7

    8.14pm: server7 done, proceeding with server8

    8.27pm: server8 done, proceeding with server9

    8.37pm: server9 done, proceeding with server10

    8.52pm: server10 done, proceeding with server12

    9.04pm: server12 done, proceeding with server11

    9.16pm: server11 done

  • Data - 17/05/2019 20:00 - 17/05/2019 21:16
  • Ultima atualização - 17/05/2019 21:16
Hits stats wrong (Resolvido)
  • Prioridade - Baixa
  • Afetando Sistema - ElasticSarch
  • Our statistics from ElasticSearch can be wrong due to a failure in data allocation, we've corrected the error on the cluster, but had to throw some log data away, because it's only temporary data we won't refill the data into the cluster, and just let it recover over the 7 days.

  • Data - 21/10/2018 22:05 - 21/10/2018 22:05
  • Ultima atualização - 21/10/2018 22:06
Disk replacement (Resolvido)
  • Prioridade - Crítico
  • Afetando Servidor - GRA5
  • We have a failing disk in server5 (GRA5), and we have to replace the disk during the evening.
    There will be downtime involved in the replacement.

    We're scheduling the replacement somewhere around 10 pm and the disk will be replaced shortly after or during the night.

    We're expecting the downtime to be roughly 30 minutes or less.
    After the replacement, we'll rebuild the raid array.

    We'll perform an additional dump of MySQL databases to our backup server prior to the replacement of the disk as a safety measure.

    We're sorry for the inconvenience caused by this - but we need to ensure the availability of the raid array.

    Update 20.15: We had a short lockup again, lasting for roughly 1 minute.

    Update 21.00: We'll request a disk replacement in a few minutes.

    Update 21.57: The server has been turned off, to get the disk replaced.

    Update 22.09: The server is back online, services are stabilizing and raid rebuild is running

  • Data - 20/09/2018 16:00 - 21/09/2018 13:00
  • Ultima atualização - 09/10/2018 13:02
Device upgrade GRA (Resolvido)
  • Prioridade - Baixa
  • Afetando Sistema - Network
  • The data center will upgrade the top of rack switches in the Gravelines data center.

    This will affect server10 (GRA10).

    The maintenance will take place starting 11 pm the 18th of September and last until 6 am the 19th.

    There will be a loss of network for up to 10 minutes.

    Server5 and server9 got completed the night between September 12 and September 13

  • Data - 18/09/2018 23:00 - 19/09/2018 06:00
  • Ultima atualização - 19/09/2018 08:04
Device upgrade RBX (Resolvido)
  • Prioridade - Baixa
  • Afetando Sistema - Network
  • The data center will upgrade the top of rack switches in the Roubaix data center.

    This will affect server6 (RBX6), server7 (RBX7) and server8 (RBX8)

    The maintenance will take place starting 11 pm the 13th of September and last until 6 am the 14th.

    There will be a loss of network for up to 10 minutes.

    Update 07.05: Maintenance completed as of 03.42 am.

  • Data - 13/09/2018 23:00 - 14/09/2018 06:00
  • Ultima atualização - 14/09/2018 07:06
Device upgrade RBX (Resolvido)
  • Prioridade - Baixa
  • Afetando Sistema - Network
  • The data center will upgrade the top of rack switches in the Roubaix data center.

    This will affect server6 (RBX6)

    The maintenance will take place starting 11 pm the 19th of September and last until 6 am the 20th.

    There will be a loss of network for up to 10 minutes.

  • Data - 19/09/2018 23:00 - 20/09/2018 06:00
  • Ultima atualização - 13/09/2018 08:00
Top of Rack switch upgrades (Resolvido)
  • Prioridade - Baixa
  • Afetando Sistema - Network
  • The data center is performing top of rack switch upgrades across RBX and GRA data centers, this means servers: server5, server6, server7, server8, server9 and server10 might be affected for up to 10 minutes, randomly during the night.

    We're sorry for the inconvenience caused by this.

    Update 07.55 am:
    top of rack switches for server5 and server9 has been updated.

    server10 is planned for the night between September 18 and September 19
    server6, server7, server8 planned for the night between September 13 and September 14

  • Data - 12/09/2018 23:00 - 19/09/2018 06:00
  • Ultima atualização - 13/09/2018 07:58
MultiPHP enabled (Resolvido)
  • Prioridade - Média
  • Afetando Servidor - GRA4
  • We'll migrate this server to a MultiPHP setup to support future versions of PHP (7.0 and 7.1)

    Currently the server runs with something called "EasyApache 3" (Provided by cPanel), we'll be upgrading to the new version called EasyApache 4 in our CloudLinux environment.

    This also means that PHP Selector will be deprecated, meaning that custom module support won't be available.

    Since this means removing old php versions (which was previously compiled from source), to a new set based on yum - it means a short downtime is expected.

    As with any other (new) server we have, we're also switching from FastCGI to mod_lsapi first of all to allow the possibility for user.ini files and php_value settings - but more importantly because also mod_lsapi isn't as buggy as FastCGI is known for.

    We've put a maintenance window of 2 hours, even though it shouldn't be needed, it should be sufficient in case any problems arise.
    We're doing our best to keep the downtime as short as possible.

    After this we'll be offering PHP version 5.6 (current version in use), 7.0 and 7.1.

    We'll enable php 5.6 on all sites after we've upgraded as the default.

    We do advise upgrading to 7.0 in case your software supports it.

    Update 9.01pm: We're starting the update in a few minutes.

    Update 9.23pm: We've completed the maintenance, we had a total downtime of 3-4 minutes meanwhile reinstalling the different versions.

    We're doing some small motifications which won't impact services.

  • Data - 21/01/2017 21:00 - 21/01/2017 21:23
  • Ultima atualização - 14/