Staff Software Engineer
About the role:
We are hiring a hands-on Staff Software Engineer to provide technical leadership for our Forecasting and Recommendations platforms, with a strong focus on production-grade AI and agentic systems.
This role centers on designing, building, and operating high-throughput, low-latency distributed systems that power forecasting, recommendations, and AI-driven decisioning at scale. You will work deeply in backend systems, infrastructure, and AI application architecture, while remaining accountable for reliability, observability, and operational excellence.
We are looking for an engineer who can contribute immediately, has shouldered real production incidents, and brings strong judgment around building stable, observable, and scalable systems, including modern agentic and LLM powered applications.
What You’ll Do :
- Design, build, and operate systems supporting forecasting, recommendations, and agentic AI workflows in production.
- Write production-quality code daily; own services end-to-end from design through on-call and incident resolution.
- Architect low-latency, high-throughput SaaS services, including APIs, data pipelines, model inference, and agent orchestration.
- Build and maintain production-grade agentic applications, including tool-using agents, workflow orchestration, and guardrails.
- Work fluently with foundational LLMs (e.g., GPT, Claude, Gemini Pro), selecting appropriate models and deployment patterns based on latency, cost, and reliability tradeoffs.
- Use frameworks and tooling such as LangChain, voice agents, and related ecosystems to accelerate development—while enforcing production discipline.
- Embrace AI-assisted development workflows (e.g., Cursor, GitHub Copilot, vibe coding paradigms) to move quickly without sacrificing quality.
- Champion observability and reliability: metrics, logging, tracing, alerting, and post-incident analysis.
- Lead and participate in production incident response, retrospectives, and systemic fixes.
- Identify architectural risks early and make design decisions that prevent outages and scalability issues.
- Reduce complexity across services, infrastructure, and processes to improve stability and team velocity.
- Provide technical guidance across teams and participate in architectural reviews beyond your immediate domain.
What We’re Looking For :
- 12+ years of professional software engineering experience building and operating production-grade distributed systems.
- A strong track record of hands-on ownership of business-critical services, including measurable improvements in latency, throughput, stability, or cost.
- Deep expertise in systems design, including service boundaries, concurrency, data modeling, failure handling, and scalability tradeoffs.
- Production experience supporting machine learning–driven systems (forecasting, recommendations, or similar), with emphasis on serving, pipelines, and infrastructure.
- Expert-level experience with AWS, including designing, deploying, and operating large-scale cloud-native systems.
- Strong hands-on experience with Kubernetes, containerized microservices, and modern CI/CD pipelines.
- Experience operating software in both on-prem data center and AWS cloud environments.
- Fluency with modern AI-assisted development tools (e.g., Cursor, GitHub Copilot) and comfort working in “vibe coding”–style workflows that favor fast iteration, tight feedback loops, and continuous refactoring.
- Proficiency in one or more backend languages commonly used for large-scale systems (e.g., Python, Java, Go, Scala).
- Bachelor’s or master’s degree in computer science, Mathematics, or a related field, or equivalent practical experience.
Staff-Level Expectations
In this role, you are expected to:
- Be a prolific, hands-on contributor who raises the bar through example.
- Own and deeply understand large portions of the codebase and system architecture.
- Simplify complex systems to increase team velocity and reliability.
- Balance technical, analytical, and product constraints to deliver pragmatic solutions.
- Set short- to medium-term (6–12 month) technical direction for your domain.
- Influence standards and architectural decisions across teams through credibility and collaboration.
- Ship multiple large services, shared libraries, or major infrastructure improvements.
Why Join This Team
- Build and operate mission-critical systems that power forecasting and recommendations at scale.
- Work in an environment that values execution, pragmatism, and modern engineering workflows.
- Operate at true Staff scope with real architectural ownership and hands-on impact.
- Solve hard systems problems with meaningful business outcomes.
Company Summary
Zeta Global is a data-powered marketing technology company with a heritage of innovation and industry leadership. Founded in 2007 by entrepreneur David A. Steinberg and John Sculley, former CEO of Apple Inc and Pepsi-Cola, the Company combines the industry’s 3rd largest proprietary data set (2.4B+ identities) with Artificial Intelligence to unlock consumer intent, personalize experiences and help our clients drive business growth.
Zeta Global is a leading AI-powered marketing technology company that enables enterprise brands to acquire, grow, and retain customers through intelligent, data-driven engagement. At the center of its innovation is the Zeta Marketing Platform (ZMP), which unifies customer data, identity, and advanced analytics to transform billions of data signals into actionable marketing intelligence and measurable business outcomes.
Publicly traded on the New York Stock Exchange (NYSE: ZETA), Zeta is redefining modern marketing with Athena by Zeta™, a superintelligent, conversational agent embedded in ZMP that personalizes the marketer’s workspace, surfaces platform generated insights through natural dialogue and recommends next-best actions to help brands accelerate and optimize their marketing outcomes.