Software Engineering
San Jose, CA, USA
About the Team
Spectro Cloud, a recognized leader in enterprise Kubernetes management with its flagship Palette platform, is seeking a hands-on Senior Software Engineer to own the resolution of complex, customer-impacting issues in production. This role is the engineering backbone behind customer success, translating field escalations into durable code fixes, backports, and long-term product improvements across commercial and federal deployments.
As a Senior Software Engineer, you will operate at the intersection of Support, Product Engineering, and Customer Reliability. You will debug deep into Kubernetes internals, Go microservices, and distributed systems; deliver patches and hotfixes on supported release branches; and drive systemic fixes that prevent recurrence. You will be a technical authority for high-severity escalations and a trusted partner to customers, TAMs, and field engineers.
About the Role
Your core mandate is to keep customer environments stable, supported, and current — while continuously improving the maintainability of the shipping product.
• Root-cause analysis and code-level fixes for customer-reported defects across the Palette platform, Kubernetes control plane and VM deployments.
• Go-based microservice debugging, patching, and backporting across multiple supported release branches.
• Kubernetes-native troubleshooting spanning CAPI, controllers, operators, CNI, CSI, and workload runtime behavior.
• Reproduction environments and test harnesses that convert customer scenarios into repeatable engineering artifacts.
• Hotfix delivery, patch releases, and CVE remediation on supported versions with strict quality and regression discipline.
• Feedback loops into Product and Engineering that convert recurring escalations into permanent product improvements.
Success in this role requires a builder-and-debugger mindset. You should be equally comfortable reading a stack trace, reading Go source, and engaging with a customer. You will use AI responsibly to accelerate triage, log analysis, and reproduction — while maintaining strong engineering accountability for every fix that ships.
Responsibilities
1. Escalation Ownership & Root Cause Analysis
• Own high-severity customer escalations end to end — from initial reproduction through code-level root cause, fix, verification, and customer confirmation.
• Debug complex Kubernetes and distributed-systems issues across control plane, cluster lifecycle, networking, storage, and workload runtime.
• Produce clear, technically rigorous RCA documents that hold up to customer, field, and executive scrutiny.
• Partner with Support, TAMs, and Customer Success to keep customers informed and unblock production impact quickly.
2. Code Fixes, Backports & Patch Delivery
• Deliver production-quality fixes in Go across the Palette codebase and Kubernetes-native components.
• Backport fixes cleanly across multiple supported release branches with strong regression discipline. • Drive hotfix, patch, and CVE releases through the Software release train, coordinating with QA, Release Engineering, and Product.
• Maintain high code-review standards on code changes — small, safe, well-tested, and well-documented.
3. Reproduction, Test Coverage & Regression Prevention
• Build reproduction environments and minimal test cases that convert one-off customer scenarios into permanent engineering assets.
• Expand unit, integration, and end-to-end test coverage to prevent regression of every fix that ships.
• Partner with QA to harden test suites against the failure modes seen in the field.
• Identify systemic gaps in observability, error handling, or upgrade paths and drive them to closure.
4. Product Feedback & Long-Term Hardening
• Identify recurring escalation patterns and drive engineering changes that eliminate their root causes. • Partner with Product Management and feature teams to feed development insights into roadmap and design reviews.
|• Improve upgrade, rollback, and day-2 operations based on real-world customer signals.
• Contribute to supportability improvements — logs, diagnostics, must-gather tooling, and self-service remediation.
5. AI-Accelerated Software Development (Responsible Innovation)
• Apply generative AI tools responsibly to accelerate log triage, stack trace analysis, reproduction scaffolding, and RCA drafting.
• Use effective prompt-engineering practices to improve the consistency and quality of AI-assisted debugging workflows.
• Validate all AI-generated artifacts before use. Increased velocity must never compromise engineering correctness, security, or compliance.
Clarity, Precision, Documentation-Driven Engineering (DDE)
• Create and maintain clear, Markdown-based RCAs, fix write-ups, and knowledge-base entries so root cause and remediation intent are documented before implementation.
• Own escalations end to end — from customer symptom through code fix, backport, verification, and follow-through on preventative work.
• Drive continuous improvement of the software development function through measurable reductions in escalation age, backport lead time, and repeat-defect rate.
• Use sustaining KPIs such as time-to-RCA, time-to-fix, backport coverage, escaped-defect rate, and customer-confirmed resolution to guide improvements.
• Collaborate effectively across Software Development, Support, Product Engineering, QA, Release Engineering, and Security teams.
Minimum Qualifications
• 3+years of hands-on software engineering experience Engineering, Escalation Engineering, or a comparable production-focused role.
• Hands-on experience with Kubernetes - operating, debugging, and modifying Kubernetes and Kubernetes-native components (controllers, operators, CAPI, CRDs) in complex production environments.
• Strong development experience - you must be comfortable reading, writing, and shipping production code, not only debugging it.
• Proven experience owning high-severity customer escalations end to end, including code-level root cause and fix delivery.
• Experience backporting fixes across multiple supported release branches with strong regression discipline.
• Experience with debugging skills across networking, storage, and control-plane behavior.
• Working knowledge of at least one major cloud platform: AWS, Azure, or GCP.
• Excellent written and verbal communication — able to hold technical authority in front of customers and executives.
Preferred Qualifications
• Go (Golang) development experience — building, debugging, and patching Go microservices and Kubernetes-native components.
• Certified Kubernetes Administrator (CKA); CKAD or CKS a plus.
• Experience with Cluster API (CAPI), controller-runtime, or writing/maintaining Kubernetes operators.
• Experience with edge, bare-metal, or virtualization platforms (VMware, KubeVirt, Hyper-V).
• Experience delivering CVE remediation and patch releases in a regulated, enterprise SaaS, or federal environment.
• Experience with MongoDB, Terraform, Ansible, or CI/CD pipelines (GitHub Actions).
• Experience using AI-assisted debugging or engineering workflows with validation and governance controls.