Skip to main content

Suggested Fixes

Overview​

Every OpsWorker investigation includes actionable recommendations organized into two categories: immediate actions to resolve the current issue, and preventive measures to stop it from recurring.

Recommendation Categories​

Immediate Actions​

Steps to fix the problem right now:

  • Specific to the current issue and environment
  • Include resource names, namespaces, and values from your cluster
  • Designed to restore service as quickly as possible
  • Typically include kubectl commands you can copy and execute

Example:

Immediate Action: Restart the deployment to recover from the OOM condition:

kubectl rollout restart deployment/api-gateway -n production

Then increase the memory limit to prevent immediate recurrence:

kubectl set resources deployment/api-gateway -n production --limits=memory=512Mi

Preventive Measures​

Longer-term fixes to address the root cause:

  • Address the underlying issue, not just the symptom
  • May involve code changes, configuration updates, or architectural adjustments
  • Help break the cycle of recurring alerts

Example:

Preventive Measure: Investigate the connection pool leak identified in the logs. The maxConnections setting in the application config should be bounded. Consider adding a connection pool metrics exporter to catch this trend earlier.

Fixing Wiring Issues​

Some of the most actionable fixes target topology wiring issues: mismatches on the connection between two resources, such as a service selector that no longer matches its pods, or a service port that does not line up with the container's target port. OpsWorker flags these during topology and configuration validation as a first-class root-cause category, and the suggested fix points directly at the misconfigured field (for example, "update the app label on the Deployment to match the Service selector").

Contextual Recommendations​

Recommendations are tailored to your specific environment:

  • Uses actual resource names, namespaces, and configuration values
  • Accounts for your cluster's topology (for example, related services that may be affected) and any wiring issues found between resources
  • Considers the specific alert type and observed failure pattern
  • References real data from the investigation (log lines, event details)
  • Is influenced by cluster-level and organization-level memory (what OpsWorker has learned about your environment) and any user-provided context from a manual, free-text investigation

Viewing Recommendations​

Recommendations appear in:

  • Slack notification: summary with key actions and commands
  • Portal investigation detail: full recommendations with context
  • Investigation chat: ask follow-up questions about any recommendation

Next Steps​