Gradient AI Serverless Inference Requests Timing Out for qwen3.5-397b-a17b
Jul 22, 19:42 UTC
Resolved - This incident has been resolved.
Jul 22, 16:50 UTC
Investigating - We are currently investigating this issue.
Jul 22, 19:42 UTC
Resolved - This incident has been resolved.
Jul 22, 16:50 UTC
Investigating - We are currently investigating this issue.
Jul 19, 15:32 UTC
Resolved - Our Engineering team has confirmed that the issue regarding volume attachment to Droplets in the NYC1, NYC3, SGP1, SYD1, and BLR1 regions has been fully resolved, and all services are now operating normally.
If you continue to experience any issues, please contact our Support team by opening a ticket. We apologize for any inconvenience caused.
Jul 19, 14:52 UTC
Monitoring - Our Engineering team has identified the root cause of the issue preventing volume attachments to Droplets in NYC1, NYC3, SGP1, SYD1, and BLR1 and has successfully implemented a fix. You should now be able to attach volumes to your Droplet(s) without encountering any further errors or issues.
We are actively monitoring the situation to ensure the fix remains effective and will provide another update once the issue has been fully resolved.
Jul 19, 13:36 UTC
Update - We are continuing to investigate this issue.
Jul 19, 13:34 UTC
Update - We are continuing to investigate this issue.
Jul 19, 09:51 UTC
Investigating - Our Engineering team is investigating an issue with attaching volume to the droplets in the NYC1, NYC3, SGP1, SYD1 and BLR1 regions. During this time you may experience an issue with attaching volume to the Droplet(s). We apologize for the inconvenience and will share an update once we have more information.
Jul 19, 09:52 UTC
Resolved - Our engineering team has resolved the outbound network connectivity issue in the BLR1 region. If you continue to experience problems, please open a ticket with our support team. We apologize for any inconvenience.
Jul 19, 09:31 UTC
Monitoring - Our engineering team has implemented the necessary fixes affecting upstream network connectivity for Droplet-based services in the BLR1 region. Users should now be able to connect to the external endpoints from Droplets in BLR1.
We are currently monitoring the situation to ensure that the service has returned to normal operation and remain stable. We appreciate your patience and will provide an update once the issue is fully confirmed as resolved.
Jul 19, 08:27 UTC
Identified - Our Engineering team has identified an issue affecting upstream network connectivity for Droplet-based services in the BLR1 region and is actively working with the relevant upstream providers to implement a resolution.
Customers may continue to experience intermittent timeouts or connectivity issues when attempting to reach external endpoints from Droplets in BLR1.
We will provide another update once the issue has been fully resolved or when additional information becomes available.
Jul 19, 06:57 UTC
Investigating - We are currently investigating an issue affecting outbound network traffic for Droplets in the BLR1 region. Customers may experience timeout errors or connectivity issues when attempting to reach external services.
Our Engineering team is actively investigating the issue and working to identify the root cause. We will provide additional updates as soon as more information becomes available.
Jul 16, 17:24 UTC
Resolved - Our Engineering team has confirmed that the underlying issue affecting Reserved IP routing in the TOR1 region has been fully resolved, and services are now operating normally.
If you continue to experience any issues, please contact our Support team by opening a ticket. We apologize for any inconvenience caused.
Jul 16, 17:01 UTC
Monitoring - Our Engineering team has implemented a fix for the issue affecting Reserved IP routing in the TOR1 region, and we are currently monitoring the results.
Inbound and outbound traffic to and from Reserved IPs in TOR1 should now be restored. If you reassigned your Reserved IP as a workaround and are still experiencing issues, please reach out to our support team.
We will continue to monitor the situation closely and provide a further update once we have confirmed the issue is fully resolved.
Jul 16, 15:30 UTC
Identified - Our Engineering team has identified the cause of the issue affecting Reserved IP routing in the TOR1 region, which was causing inbound and outbound traffic to and from Reserved IPs in TOR1 to fail.
Detaching and reassigning the Reserved IP to the Droplet may restore connectivity in the meantime. We recommend trying this workaround if you are still experiencing issues with your Reserved IP in TOR1.
Our engineering team is now working on implementing a fix. We will provide updates as more information becomes available.
Jul 16, 13:39 UTC
Investigating - Our Engineering team is investigating an issue affecting Reserved IP routing in the TOR1 region. The issue is causing inbound and outbound traffic to and from Reserved IPs in TOR1 to fail.
In the meantime, detaching and reassigning the Reserved IP to the Droplet may restore connectivity. We recommend trying this workaround if you are experiencing issues with your Reserved IP in TOR1.
Our engineering team is actively investigating the root cause and working towards a resolution. We will provide updates as more information becomes available.
Jul 15, 21:46 UTC
Resolved - The deployed fix has successfully restored full functionality, and our monitoring shows that system performance has completely stabilized. Response times for all Gemma 4 inference workflows have returned to normal baseline levels.
We will continue to track platform stability moving forward to ensure long-term reliability. We apologize for any disruption this may have caused to your workflows and appreciate your patience throughout the recovery process.
Jul 15, 10:06 UTC
Monitoring - A fix has been deployed to resolve the issue. We are closely monitoring system performance to ensure full recovery and normal response times for all Gemma 4 inference workflows.
Jul 14, 11:07 UTC
Identified - We are currently experiencing an issue affecting customers using the Gemma 4 model on our Serverless Inference platform. Customers may experience significantly increased latency or request timeouts.
Our Engineering team has identified a backend configuration issue as the root cause, which is temporarily impacting model performance. Please be assured that our Engineering team is actively working on a fix and is treating this issue with high priority.
We sincerely apologise for any inconvenience this may have caused and appreciate your patience and understanding. If you have any further questions, please create a support ticket so that we can investigate your specific case further.
Jul 13, 22:24 UTC
Monitoring - A fix has been deployed to resolve the backend configuration issue. We are closely monitoring system performance to ensure full recovery and normal response times for all Gemma 4 inference workflows.
Jul 13, 19:04 UTC
Identified - We are currently experiencing an issue where customers using the Gemma 4 model on our Serverless Inference and Dedicated Inference platforms may experience severe latency or request timeouts.
Our engineering team has identified a backend configuration issue as the root cause, which is temporarily degrading performance. We are actively working on a fix to restore normal response times and will provide another update as soon as the mitigation is in place
Jul 11, 01:27 UTC
Resolved - The issue affecting Kubernetes deployments in the NYC1 region has been resolved. Our investigation found intermittent DNS timeouts affecting a small number of DOKS clusters, with affected worker nodes running on shared-CPU Droplets. This is a documented limitation for latency-sensitive cluster DNS workloads such as CoreDNS.
The affected clusters are currently functional. To reduce the risk of recurrence, we recommend running CoreDNS on non-shared/dedicated CPU node pools and using sufficient CoreDNS replicas.
If you continue to see DNS failures or NodeNotReady events, please open a support ticket so we can investigate that cluster specifically.
Jul 10, 01:18 UTC
Monitoring - The issue affecting Kubernetes deployments in NYC1 has subsided. Workloads should now be functioning normally.
Our Engineering team is continuing to monitor the affected systems to confirm full resolution. We'll update this page if any further action is needed. We apologize for the inconvenience this may have caused.
Jul 9, 20:31 UTC
Investigating - Our Engineering team is investigating an issue affecting Kubernetes deployments in NYC1. Users may see intermittent DNS failures and NodeNotReady events from application workloads during this time.
We apologize for the inconvenience, we'll share new information on this page as soon as it is available.
Jul 5, 00:53 UTC
Resolved - Our Engineering team has confirmed that the issue with the Deepseek V4 Pro model in Serverless Inference and Agent Platform has been fully resolved. The model is now operational, and users should be able to use it without experiencing any errors.
If you continue to experience any problems, please open a ticket with our Support team. We apologize for any inconvenience this may have caused and appreciate your patience.
Jul 4, 23:18 UTC
Monitoring - Our Engineering team has implemented a fix for the issue with the Deepseek V4 Pro model in Serverless Inference and Agent Platform. The model should now be operational, and users should no longer receive error 429 messages when attempting to use it. We are currently monitoring the situation to ensure the fix is successful and the model is functioning as expected. We will post an update if any further issues arise. If you continue to experience problems, please open a ticket with our Support team. We apologize for any inconvenience this may have caused.
Jul 4, 22:13 UTC
Investigating - Our Engineering team is investigating reports of an incident affecting the Deepseek V4 Pro model in Serverless Inference and Agent Platform. Users may experience errors when attempting to use this model, specifically receiving error 429 messages. We apologize for the inconvenience and are working to resolve the issue as soon as possible. We will provide an update once we have more information.
Jul 4, 15:10 UTC
Resolved - Our Engineering team has confirmed that the issue impacting Droplets in the NYC3 region has been fully resolved. Users can now perform actions on their Droplets or create new Droplets in this region without any issues.
We appreciate your patience while we worked to resolve this issue. If you continue to experience any problems, please open a support ticket from your account so our team can investigate further.
Jul 4, 14:42 UTC
Monitoring - Our Engineering team has implemented a fix to address the issue affecting Droplets in the NYC3 region. The issue, which began at 12:38 UTC, caused disruptions to customers attempting to perform actions on their Droplets in the NYC3 region.
During this time, customers may also have experienced errors when trying to create Droplets in this region.
We are actively monitoring the situation to ensure the fix remains effective and will provide another update once the issue has been fully resolved.
Jul 7, 15:30 UTC
Completed - The scheduled maintenance has been completed.
Jul 7, 10:30 UTC
In progress - Scheduled maintenance is currently in progress. We will provide updates as necessary.
Jul 4, 09:05 UTC
Scheduled - Start: 2026-07-07 10:30 UTC
End: 2026-07-07 15:30 UTC
During the above maintenance window, our Networking Engineering team will be making changes to the core networking infrastructure on Baremetal GPU to enhance performance and scalability in the NYC2 region.
Expected impact:
We do not expect any customer impact during this maintenance. However, if any unexpected issues arise, our team will continuously monitor the situation and take appropriate corrective actions, including rolling back the change if necessary to restore service stability. We will make every effort to minimize any impact.
If you have any questions related to the maintenance, please don’t hesitate to reach out to us through https://cloudsupport.digitalocean.com/s/createticket
Jul 3, 08:20 UTC
Resolved - Our engineering team has resolved the issue with Managed Database Clusters. All database systems should now be operating normally. If you continue to experience problems, please open a ticket with our support team. We apologize for any inconvenience.
Jul 3, 06:08 UTC
Monitoring - Our engineering team has implemented a fix to resolve the issue with Managed Database Clusters and is monitoring the situation. We will post an update as soon as the issue is fully resolved.
Jul 3, 02:28 UTC
Update - Our engineering team is still investigating the issue affecting Managed Database Clusters. Users may still encounter delays creating/scaling/forking/restoring the aforementioned Managed Database Clusters.
We are working to resolve this as soon as possible. we apologize for the inconvenience. We will post an update here once we have more information.
Jul 2, 22:04 UTC
Investigating - Our engineering team is investigating an issue affecting Managed Database Clusters. Currently, users may encounter delays when creating/scaling/forking and restoring Standard MySQL, Standard PostgreSQL, OpenSearch, Kafka, and Valkey clusters through the Cloud Control Panel or API.
We apologize for the inconvenience and will provide further updates as soon as more information is available.
Jul 1, 16:28 UTC
Resolved - The intermittent DNS lookup failures between Managed Kubernetes and Managed Database hostnames in the FRA1 region have been resolved for now but please let us know if you see the issue again. Connectivity has remained stable, and all systems are operating normally. If you continue to experience any problems, please open a ticket with our support team. Thank you for your patience, and we apologize for any inconvenience.
Jul 1, 12:15 UTC
Update - We are continuing to investigate this issue.
Jul 1, 11:35 UTC
Investigating - As of 06:39 UTC, our Engineering team is investigating reports of intermittent DNS lookup failures for Managed Database hostnames from Managed Kubernetes, primarily affecting connections from Managed Kubernetes to Managed Database hostnames.
At this point, customers in the FRA1 region may experience intermittent connectivity issues.
We apologize for the inconvenience and will share an update once we have more information.
Jul 1, 13:28 UTC
Resolved - Our Engineering team has confirmed that the issue impacting Droplet resizes using the "Downscale Anytime" option has been fully resolved. Users can now resize their Droplets using the "Downscale Anytime" option without any issues.
We appreciate your patience while we worked to resolve this issue. If you continue to experience any problems, please open a support ticket from your account so our team can investigate further.
Jul 1, 12:43 UTC
Monitoring - Our Engineering team has identified the root cause of the issue affecting Droplet resizes using the "Downscale Anytime" option and has implemented a fix. Users should now be able to resize their Droplets using the "Downscale Anytime" option without experiencing any issues or errors.
We are actively monitoring the situation to ensure the fix remains effective and will provide another update once the issue has been fully resolved.
Jul 1, 11:13 UTC
Investigating - Our Engineering team is investigating an issue affecting Droplet resizes using the "Downscale Anytime" option. At this time users may find this option unavailable or unresponsive in the Cloud Control Panel.
As a workaround, resizes can be performed via the API while we work to resolve this issue.
We apologize for the inconvenience and will share an update once we have more information.
Jul 2, 03:38 UTC
Resolved - This incident has been resolved, and our teams continue to monitor the results. If you experience any further issues, please contact support.
Jul 1, 08:39 UTC
Identified - Our Engineering team has identified the issue causing Agents to experience timeouts while retrieving data from Knowledge Bases.
Please be assured that our Engineering team is actively working on a fix and is treating this issue with high priority.
We sincerely apologise for any inconvenience this may have caused and appreciate your patience and understanding. If you have any further questions, please create a support ticket so that we can investigate your specific case further.
Jul 1, 06:25 UTC
Investigating - Our Engineering team is currently investigating an issue where Agents are experiencing timeouts while retrieving data from Knowledge Bases. As a result, affected Agents may fail to retrieve data and return the following error:
"Failed to retrieve data from Knowledge base(s) - timeout"
Please be assured that we are treating this as a high-priority issue and are actively working to mitigate it.
We sincerely apologise for any inconvenience this may have caused and appreciate your patience and understanding. If you have any further questions, please create a support ticket so that we can investigate your specific case further.
Jun 27, 18:24 UTC
Resolved - Between 00:00 UTC & 16:00 UTC today, our Engineering team identified an issue affecting backup operations on Droplets. During this period, backups scheduled within this window may not have been created and may appear as missing.
Our team has taken necessary measures to resolve the issue, and we can confirm that the backup service has been restored and is now functioning normally. Upcoming scheduled backups should be performed as expected.
We sincerely apologize for any inconvenience this may have caused and appreciate your understanding. However, if you have any further questions or concerns, please create a support ticket for further analysis.
Jun 27, 05:15 UTC
Resolved - Our Engineering team has resolved the issue that was causing HTTP 400 errors for requests to Anthropic models. Users should now be able to access Anthropic models without any issues.
If you continue to experience problems, please open a ticket with our Support team so we can investigate further.
We apologize for any inconvenience this may have caused.
Jun 27, 04:43 UTC
Monitoring - Our Engineering team has mitigated an issue with Anthropic models. Previously, users may have encountered 400 errors when attempting to use any Anthropic model.
Although the root cause of the issue is still being addressed by Anthropic, users should now be able to access and use Anthropic models again. We will continue to monitor the situation and provide updates if necessary. If you continue to experience problems, please open a ticket with our Support team. We apologize for any inconvenience this may have caused.
Jun 24, 12:30 UTC
Completed - The scheduled maintenance has been completed.
Jun 24, 09:30 UTC
In progress - Scheduled maintenance is currently in progress. We will provide updates as necessary.
Jun 22, 09:40 UTC
Scheduled - Start: 2026-06-24 9:30 UTC
End: 2026-06-24 12:30 UTC
During the above window, our Networking team will be making changes to the core networking infrastructure to improve performance and scalability in the SFO1 region.
Expected impact:
These changes are designed and tested to be seamless. We do not expect any customer impact during the mentioned timeframe. If an unexpected issue arises, there could be a temporary loss of connectivity for Droplets and its dependent services, such as Managed Databases, Load Balancers, App Platform, and Managed Kubernetes, in the SFO1 region. We will endeavor to minimize any such impact.
If you have any questions related to this issue, please send us a ticket from your cloud support page. https://cloudsupport.digitalocean.com/s/createticket
Jun 19, 16:37 UTC
Resolved - Between 13:22 & 14:26 UTC today, users have experienced missing monitoring graphs for Droplets (with DO agent installed), Load Balancers, Databases, and other services within the Cloud Control Panel.
Our engineering team has identified the root cause of the issue and has taken appropriate steps to restore functionality. We can confirm that services have been restored and are functioning as expected.
We apologize for any inconvenience this may have caused. If you continue to experience issues viewing monitoring graphs, please create a support ticket for further analysis. Thank you for your patience and understanding
Jun 19, 14:50 UTC
Monitoring - Our engineering team has implemented the necessary fixes to address the issue affecting the visibility of monitoring graphs within the Cloud Panel. Users should now be able to view monitoring graphs for their services, including Droplets with DO Agent installed, Load Balancers, Databases, etc.
We are currently monitoring the situation to ensure that the service has returned to normal operation and remain stable. We appreciate your patience and will provide an update once the issue is fully confirmed as resolved.
Jun 19, 14:00 UTC
Investigating - Our Engineering team is currently investigating an issue affecting the visibility of monitoring graphs within the Cloud Panel. During this period, users may notice missing or unavailable monitoring graphs for services such as Droplets(with DO agent installed), Load Balancers, Databases, etc.
We apologize for the inconvenience caused. We'll update once we have more information
Jun 17, 18:00 UTC
Completed - The scheduled maintenance has been completed.
Jun 17, 09:00 UTC
In progress - Scheduled maintenance is currently in progress. We will provide updates as necessary.
Jun 17, 05:53 UTC
Scheduled - Core Infrastructure Maintenance
Start: 2025-06-17 09:00 UTC
End: 2025-06-17 18:00 UTC
During the above time, our Engineering Team will be performing maintenance to failover some internal databases from one cluster to another.
Existing infrastructure, including Droplets and Droplet-based services, should continue running without issue. There is no network disruption to existing services expected as part of this maintenance. However, there are dependencies on multiple services. During the failover, there may be customer impacts that should be brief and transitory.
Multiple teams will be engaged to keep downtime to a minimum and mitigate any impact that does occur. We’ll post updates here for any unexpected changes to this scheduled maintenance, as well as progress updates during the maintenance itself.
If you have any questions related to this issue please send us a ticket from your cloud support page. https://cloudsupport.digitalocean.com/s/createticket
Thank you,
Team DigitalOcean
Jun 17, 00:29 UTC
Resolved - The connectivity issues affecting our Serverless/GenAI Inference API have been fully resolved.
Our engineering teams successfully completed the connectivity restoration. All systems should be functioning normally, and the endpoint should be fully operational.
Jun 16, 23:29 UTC
Monitoring - Our engineering teams have successfully begun implementing mitigation steps to resolve the connectivity issues affecting the inference API. We will provide another update once the API has fully recovered and error rates return to normal.
Jun 16, 22:00 UTC
Investigating - We are actively investigating an issue causing elevated HTTP 500 error rates for customers utilizing our Serverless/GenAI Inference API.
Customer Impact: Customers making calls to the inference API—specifically targeting /v1/* endpoints—will experience intermittent HTTP 500 errors and failed requests.
Jun 18, 17:30 UTC
Completed - The scheduled maintenance has been completed.
Jun 18, 14:30 UTC
In progress - Scheduled maintenance is currently in progress. We will provide updates as necessary.
Jun 16, 14:32 UTC
Scheduled - Start: 2026-06-18 14:30 UTC
End: 2026-06-18 17:30 UTC
During the above window, our Networking team will be making changes to the core networking infrastructure to improve performance and scalability in the SGP1 region.
Expected impact:
We do not anticipate any downtime for Droplets or Droplet-related services, including Managed Databases, Load Balancers, App Platform, and Managed Kubernetes, as this maintenance has been carefully designed and tested to be seamless. In the unlikely event that an undetected misconfiguration occurs, a subset of customers could experience temporary network disruption. We will endeavor to keep this to a minimum for the duration of the change.
If you have any questions related to this issue, please send us a ticket from your cloud support page. https://cloudsupport.digitalocean.com/s/createticket