Reading view

MongoDB Maintenance - BLR1, NYC3, SFO2, SGP1, SYD1, TOR1

Apr 9, 22:51 UTC
Completed - The scheduled maintenance has been completed.

Apr 9, 18:00 UTC
In progress - Scheduled maintenance is currently in progress. We will provide updates as necessary.

Apr 7, 18:23 UTC
Update - We will be undergoing scheduled maintenance during this time.

Apr 7, 18:18 UTC
Scheduled - Start: 2026-04-09 18:00 UTC
End: 2026-04-10 24:00 UTC

During the above window, our Engineering team will perform maintenance on core MongoDB services in the BLR1, NYC3, SFO2, SGP1, SYD1 & TOR1regions to enhance security and improve auditing and compliance. Please note that existing databases and workloads will continue to function normally and will not be impacted.

Expected Impact:

We do not anticipate any service disruptions during this window. Your existing databases and workloads will continue to run normally without interruption.

In the event that an unexpected issue occurs, administrative actions, such as creating, deleting, or scaling Managed MongoDB databases in the BLR1, NYC3, SFO2, SGP1, SYD1 & TOR1 regions, may experience delays.

If an unexpected issue arises, we will work to keep any impact to a minimum and may revert the changes if required.

If you have any questions related to this event, please open a ticket from your cloud support page: https://cloudsupport.digitalocean.com/s/createticket

  •  

Control Plane

Apr 7, 17:49 UTC
Resolved - Our Engineering team has resolved the control plane disruption that occurred from 17:06 to 17:18 UTC. During this time, users may have experienced intermittent issues with managing their resources through the Cloud Control Panel or DigitalOcean API. The root cause of the disruption was identified and addressed, and all services are now operating normally.

If you continue to experience any problems, please open a ticket with our Support team. We apologize for any inconvenience this may have caused.

  •  

Serverless Inference - High error rates for open source models ( Qwen 3 32B)

Apr 7, 15:50 UTC
Resolved - Service has been fully restored, and the model is now operating normally. We have implemented improvements to enhance stability and reduce the likelihood of similar issues in the future.

Apr 7, 12:55 UTC
Identified - We are currently investigating reports of elevated latency affecting requests to this model when using Serverless Inference and Agents.

Earlier observations indicated increased error rates for the open-source Qwen 3 32B model. The Ray dashboard also showed multiple workers in a pending state, suggesting capacity constraints.

Our analysis determined that the model was experiencing higher-than-expected request volume without sufficient resources to scale accordingly. To address this, the node pool size has been increased to improve available capacity. However, there are still insufficient nodes to fully support the desired number of model replicas.

Following the node pool expansion, a new pod-related error has been identified. Our Engineering team is actively working to resolve this issue and restore full service performance.

Apr 7, 12:49 UTC
Investigating - Serverless inference for alibaba-qwen3-32b (Qwen 3 32B) in tor1 is experiencing high error rates starting at 10:46 UTC.

  •  

Serverless Inference Issue

Apr 6, 18:02 UTC
Resolved - This incident has been resolved.

Apr 6, 15:15 UTC
Monitoring - A fix has been implemented and we are monitoring the results.

Apr 6, 12:28 UTC
Investigating - Our Engineering team is investigating an issue with Serverless inference.

At this time, users may experience high error rates for open source models (llama 3.3 70b).

We apologize for the inconvenience and will share an update once we have more information.

  •  

MongoDB Maintenance - AMS3, ATL1, LON1, NYC1, NYC2, SFO3

Apr 7, 22:25 UTC
Completed - The scheduled maintenance has been completed.

Apr 7, 18:00 UTC
In progress - Scheduled maintenance is currently in progress. We will provide updates as necessary.

Apr 5, 18:12 UTC
Scheduled - Start: 2026-04-07 18:00 UTC
End: 2026-04-08 00:00 UTC

During the above window, our Engineering team will perform maintenance on core MongoDB services in the AMS3, ATL1, LON1, NYC1, NYC2 & SFO3 regions to enhance security and improve auditing and compliance. Please note that existing databases and workloads will continue to function normally and will not be impacted.

Expected Impact:

We do not anticipate any service disruptions during this window. Your existing databases and workloads will continue to run normally without interruption.

In the event that an unexpected issue occurs, administrative actions, such as creating, deleting, or scaling Managed MongoDB databases in the AMS3, ATL1, LON1, NYC1, NYC2 & SFO3 regions, may experience delays.

If an unexpected issue arises, we will work to keep any impact to a minimum and may revert the changes if required.

If you have any questions related to this event, please open a ticket from your cloud support page: https://cloudsupport.digitalocean.com/s/createticket

  •  

FRA1 MongoDB Maintenance

Mar 27, 23:16 UTC
Completed - The scheduled maintenance has been completed.

Mar 27, 19:00 UTC
In progress - Scheduled maintenance is currently in progress. We will provide updates as necessary.

Mar 25, 19:01 UTC
Scheduled - Start: 2026-03-27 19:00 UTC
End: 2026-03-28 02:00 UTC

During the above window, our Engineering team will perform maintenance on core MongoDB services in the FRA1 region to enhance security and improve auditing and compliance. Please note that existing databases and workloads will continue to function normally and will not be impacted.

Expected Impact:

We do not anticipate any service disruptions during this window. Your existing databases and workloads will continue to run normally without interruption.

In the event that an unexpected issue occurs, administrative actions, such as creating, deleting, or scaling Managed MongoDB databases in the FRA1 region, may experience delays.

If an unexpected issue arises, we will work to keep any impact to a minimum and may revert the changes if required.

If you have any questions related to this event, please open a ticket from your cloud support page: https://cloudsupport.digitalocean.com/s/createticket

  •  

App platform seeing delays in deployments across FRA1 region

Mar 20, 12:01 UTC
Resolved - The issue impacting delays in App Platform deployments has been confirmed to be resolved. Between approximately 00:08am UTC & 11:46am UTC, users may have noticed delays while creating or updating apps, or may have encountered failed deployments. For failed deployments, please trigger a redeploy, which should successfully resolve the issue.

We confirmed that the service is functioning as expected. Once again, we sincerely apologize for the inconvenience caused and appreciate your understanding.

However, if you continue to experience any issues, please don't hesitate to raise a support ticket for further investigation. We'll be happy to assist you.

Mar 20, 11:14 UTC
Monitoring - Our Engineering team has deployed a fix to resolve the issue impacting new App Platform deployments using Dedicated Egress IP in FRA1 region. We are actively monitoring the situation to ensure stability and will provide an update once the incident has been fully resolved. Thank you for your patience and we apologize for the inconvenience.

Mar 20, 10:32 UTC
Investigating - Our engineers are currently investigating an issue impacting new App Platform deployments using Dedicated Egress IP in FRA1 region. During this time, some users may experience delay when creating new App Platform apps or deploying existing apps. Existing apps are not affected and should continue to function normally. We apologize for any inconvenience, and we'll share more information as it becomes available.

  •  

Gradient AI Platform agents and services Accessibility

Mar 20, 14:14 UTC
Resolved - Our Engineering team has implemented a fix, the issues impacting Gradient AI Platform have been resolved. All agents are back up and healthy. Service has been fully restored.

Mar 20, 14:04 UTC
Monitoring - A fix has been implemented and services have been restored. We are continuing to monitor the system to ensure stability. We will provide further updates if needed.

Mar 20, 11:05 UTC
Update - We've identified the issue and are actively working to restore the affected services. We're making steady progress and closely monitoring the situation. Further updates will be shared as they become available.

Mar 20, 09:51 UTC
Identified - We’ve identified the issue and are currently working on restoring the services. We’ll continue to provide updates as progress is made.

Mar 20, 08:50 UTC
Investigating - We are currently investigating issue affecting the accessibility of agents and services on the Gradient AI Platform. Users may experience failures or unresponsiveness when attempting to use these features. Our engineering team is actively working to identify the root cause and restore full functionality. We apologize for the inconvenience and will share an update once we have more information.

  •  

Gradient AI model availability

Mar 17, 19:49 UTC
Resolved - Our Engineering team has implemented a fix, the issues impacting model availability and performance have been resolved. All models, including those previously degraded, are back up and healthy. Service has been fully restored.

Mar 17, 15:00 UTC
Investigating - Our Engineering team is investigating reports of Gradient AI model availability issues impacting multiple models. Users may experience issues with models availability, including Llama3.1-8b and Qwen3-32b, as well as embedding models such as GTE Large (v1.5), All-MiniLM-L6-v2, Multi-QA-mpnet-base-dot-v1, and Qwen3 Embedding 0.6B.

Additionally, Guardrails are not available, affecting associated agents, and users attempting to run inference on the Llama3.3-70b model will see degraded performance.

We apologize for the inconvenience and will share an update once we have more information.

  •  

Cloud Control Panel and API

Mar 16, 17:39 UTC
Resolved - From 16:14 to 16:38 UTC, Our Engineering team observed an issue impacting Cloud control panel and API. During this time, users may experienced errors when trying to access the Cloud control panel and when trying to use the API. Our team has fully resolved the issues as of 16:38 UTC. If you continue to experience problems, please open a ticket with our support team from within your Cloud Control Panel. We apologize for any inconvenience caused.

  •  

Degraded performance with BYOK Anthropic models

Mar 15, 03:31 UTC
Resolved - The issue is now resolved, all Anthropic BYOK models in Gradient AI should work normally.
Contact support if issues persist.

Mar 15, 02:55 UTC
Investigating - Our Engineering team is investigating an issue related to all Gradient AI agents and serverless inference that require BYOK Anthropic modles.
Impacted users may experience degraded performance.
We will provide an update as soon as possible

  •  

Delay in App Platform Deployments

Mar 14, 01:47 UTC
Resolved - As of 23:00 UTC, our Engineering team has confirmed that the issue causing delays in App Platform deployments has been fully resolved. The fix implemented earlier has been successful, and we are no longer seeing any delays or errors with deployments.

Users should now be able to deploy their apps successfully and without any issues. We apologize again for the inconvenience caused.

However, if you continue to experience any issues, please don't hesitate to raise a support ticket for further investigation.

Mar 13, 23:39 UTC
Monitoring - After working with our upstream provider, our Engineering team has implemented a fix to resolve the issue that was causing delays in the deployment of new apps, and they are currently monitoring the situation.

During this time, users should no longer experience issues with creating new apps and all the stalled creation events should provision completely.

We will post an update as soon as the issue is fully resolved.

Mar 13, 22:01 UTC
Identified - Our Engineering team is starting to see delays once again with new App Platform deployments. During this time, users may still experience delays with deploying new apps. We're working with our upstream provider to resolve the issue.

We again apologize for the inconvenience. We will post further updates once we have more information.

Mar 13, 21:30 UTC
Monitoring - Starting at 20:40 UTC, users may have seen delays with deploying new apps on App Platform.

At this time, our Engineering team is seeing signs of recovery, and users should be able to deploy new apps without issue. We're currently monitoring the situation to ensure full recovery.

We apologize for the inconvenience. We'll post an update once the issue has been confirmed to be resolved.

  •  

Newly Created Managed Kubernetes Nodes

Mar 13, 16:35 UTC
Resolved - Our Engineering team has confirmed the resolution of the issue impacting DNS timeouts for newly provisioned Managed Kubernetes nodes. At this time all cluster services should now be functioning normally. If you continue to experience problems, please open a ticket with our support team. We apologize for any inconvenience.

Mar 13, 13:55 UTC
Monitoring - Our Engineering team has implemented a fix to address the issue causing DNS timeouts for newly provisioned Managed Kubernetes nodes. Further investigation has confirmed that this issue primarily affected customers utilizing a NAT Gateway within their VPC and running a VPC-native cluster. We are actively monitoring the situation to ensure overall stability.

We appreciate your patience and will provide a further update once the issue is fully confirmed to be resolved.

Mar 13, 12:32 UTC
Identified - Our Engineering team is investigating an issue impacting newly provisioned Managed Kubernetes nodes. During this time, Only customers who run a NAT Gateway in their VPC and a VPC-native clusters are affected and may experience DNS timeouts. We apologize for the inconvenience and will share an update once we have more information.

Mar 13, 11:26 UTC
Investigating - Our Engineering team is investigating an issue impacting newly provisioned Managed Kubernetes nodes. During this time, new nodes may experience DNS timeouts, which could temporarily affect cluster services. We apologize for the inconvenience and will share an update once we have more information.

  •  

Ubuntu/Debian Package Mirror Failure

Mar 9, 19:23 UTC
Resolved - From 17:50 to 19:06 UTC, Our Engineering team observed an issue with mirrors.digitalocean.com. During this time, users may have experienced errors when trying to update packages on Debian and Ubuntu Images. Our team has fully resolved the issues as of 19:06 UTC. If you continue to experience problems, please open a ticket with our support team from within your Cloud Control Panel. We apologize for any inconvenience caused.

  •  

HTTP 522 Error on App Platform

Mar 6, 21:22 UTC
Resolved - Our Engineering team identified an issue affecting the App Platform. During the incident, users may have experienced HTTP 522 (Connection Timed Out) errors when accessing their apps. The issue seems to be resolved now.

We apologize for the inconvenience caused. If you continue to experience any related errors, please contact our Support team by opening a ticket at https://www.digitalocean.com/support/contact/.

  •  

App Platform Deployments

Mar 6, 01:12 UTC
Resolved - As of 00:22 UTC, our Engineering team has confirmed that the issue causing delays in App Platform deployments has been fully resolved. The fix implemented earlier has been successful, and we are no longer seeing any delays or errors with deployments.

Users should now be able to deploy their apps successfully and without any issues. We apologize again for the inconvenience caused.

However, if you continue to experience any issues, please don't hesitate to raise a support ticket for further investigation.

Mar 6, 00:32 UTC
Monitoring - Our Engineering team has implemented a fix to address the issue causing delays in App Platform deployments. We are actively monitoring the situation to ensure overall stability.

We appreciate your patience and will provide a further update once the issue is fully confirmed to be resolved.

Mar 5, 23:28 UTC
Investigating - Our Engineering team is currently investigating an issue impacting App Platform deployments. During this time, users may experience a delay or failure when deploying new and existing App Platform apps.

We apologize for any inconvenience, and we'll share more information as it becomes available.

  •  

Internal Load Balancers Connectivity

Mar 5, 01:52 UTC
Resolved - From 19:57 UTC to 01:03 UTC, customers may have experienced connectivity issues between Internal Load Balancers and their associated target droplets, which could have resulted in service disruption or traffic routing failures.

Our Engineering team has confirmed full resolution of the issue, and Internal Load Balancers should now be functioning normally.

If you continue to experience any problems, please open a ticket with our Support team. We apologize for any inconvenience caused.

Mar 5, 01:19 UTC
Monitoring - Our Engineering team has implemented mitigation measures to address the connectivity issues affecting Internal Load Balancers and their associated target droplets. We are actively monitoring the situation to ensure stability and to prevent any recurrence.

We will provide a further update once we confirm the issue is fully resolved.

Mar 5, 00:23 UTC
Investigating - Our Engineering team is investigating an issue affecting Internal Load Balancers. Customers may experience connectivity loss between Internal Load Balancers and their associated target droplets.

We apologize for the inconvenience and will share an update as soon as more information becomes available.

  •  

Core Infrastructure Maintenance in All Regions 2026-03-03 10:00 UTC

Mar 6, 15:55 UTC
Completed - This scheduled maintenance is now complete across all regions. Thank you for your patience and understanding throughout this process.

Mar 6, 11:00 UTC
In progress - Scheduled maintenance is currently in progress. We will provide updates as necessary.

Mar 5, 13:05 UTC
Scheduled - Phase 2 maintenance is complete. Phase 3 is scheduled to begin at March 06, 11:00 UTC.

Mar 5, 10:00 UTC
In progress - Scheduled maintenance is currently in progress. We will provide updates as necessary.

Mar 3, 13:03 UTC
Scheduled - Phase 1 maintenance is complete. Phase 2 is scheduled to begin at March 05, 10:00 UTC.

Mar 3, 10:00 UTC
In progress - Scheduled maintenance is currently in progress. We will provide updates as necessary.

Mar 1, 10:44 UTC
Scheduled - Start: 2026-03-03 10:00 UTC
End: 2026-03-06 13:00 UTC

Hello,

During the above window, our Engineering team will be performing maintenance on core control plane infrastructure in all regions. Please note that the existing infrastructure will continue running without issue.This maintenance will be carried out in three phases as outlined below:

March 03, 10:00 to 13:00 UTC
March 05, 10:00 to 13:00 UTC
March 06, 11:00 to 13:00 UTC


We do not anticipate any impact; however, there is a small possibility that control panel functionality specifically CRUD (Create, Read, Update, Delete) operations may be affected during the maintenance window. All running workloads are expected to continue operating normally without interruption.

Our team will be actively monitoring the environment throughout the maintenance, and any unexpected events will be promptly communicated through our status page.

If you have any questions or concerns regarding this maintenance, please feel free to open a support ticket from within your account. We’re here to help.

  •  

Core Infrastructure Maintenance in SFO2 and SFO3

Mar 2, 16:00 UTC
Completed - The scheduled maintenance has been completed.

Mar 2, 13:00 UTC
In progress - Scheduled maintenance is currently in progress. We will provide updates as necessary.

Feb 28, 13:20 UTC
Scheduled - Start: 2026-03-02 13:00 UTC
End: 2026-03-02 16:00 UTC

During the above window, our Engineering team will be performing maintenance on core control plane infrastructure in SFO2 and SFO3. Please note that the existing infrastructure will continue running without issue.

We do not anticipate any impact; however, there is a small possibility that control panel functionality specifically CRUD (Create, Read, Update, Delete) operations may be affected during the maintenance window. All running workloads are expected to continue operating normally without interruption.

Our team will be actively monitoring the environment throughout the maintenance, and any unexpected events will be promptly communicated through our status page.

If you have any questions or concerns regarding this maintenance, please feel free to open a support ticket from within your account. We’re here to help.

  •  

Intermittent Errors with Llama 3.3-70B

Feb 26, 22:23 UTC
Resolved - Issue resolved.
Cause: A few requests made to the Llama 3.3-70B model caused issues.
Impact: Intermittent errors when interacting with the model through serverless inference and/or with agents created using this model.
Contact support if issues persist.

Feb 26, 21:52 UTC
Monitoring - Fix deployed. Monitoring resources related to the Llama 3.3-70B.
Users should no longer experience intermittent errors when making serverless inference requests via APIs and Agents . Awaiting confirmation before closure.

Feb 26, 16:00 UTC
Investigating - We are currently investigating an issue affecting the Llama 3.3-70B model.
Symptoms: Users may encounter intermittent errors when making serverless inference requests via APIs and Agents.
Current Status: Our engineering team is actively investigating the issue to determine the root cause.

  •  
❌