JobBobsReal-time global job discoveryLive

Site Reliability Engineer 3

Granicus · United States · 2026-07-29

seniorRemote
Apply on the employer's site

About this role

The Company
Serving the People Who Serve the People
Granicus is driven by the excitement of building, implementing, and maintaining technology that is transforming the Govtech industry by bringing governments and its constituents together. We are on a mission to support our customers with meeting the needs of their communities and implementing our technology in ways that are equitable and inclusive. Granicus has consistently appeared on the GovTech 100 list over the past 5 years and has been recognized as the best companies to work on BuiltIn.
Over the last 25 years, we have served 5,500 federal, state, and local government agencies and more than 300 million citizen subscribers power an unmatched Subscriber Network that use our digital solutions to make the world a better place. With comprehensive cloud-based solutions for communications, government website design, meeting and agenda management software, records management, and digital services, Granicus empowers stronger relationships between government and residents across the U.S., U.K., Australia, New Zealand, and Canada. By simplifying interactions with residents, while disseminating critical information, Granicus brings governments closer to the people they serve—driving meaningful change for communities around the globe.
Want to know more? See more of what we do here.
Job Summary
Granicus is the leading provider of citizen engagement technologies and services for the public sector, bringing governments closer to the people they serve with the first-and-only Civic Engagement Platform. Granicus works with more than 5,500 government organizations and connects more than 280 million people in the largest Citizen Subscriber Network of its kind.
What Your Impact Will Look Like
Granicus is seeking a Site Reliability Engineer (SRE3) with strong AIOps capabilities to modernize reliability engineering through observability, automation, and AI-assisted operations. In this role, you will improve service reliability, reduce operational toil, accelerate incident response, and help build scalable, resilient platforms supporting both traditional and AI/ML-powered workloads.
What your impact will look like
AI, MCP & AIOps
Lead adoption of AI-first SRE practices across monitoring, incident response, and automation
Design and implement MCP-based integrations connecting systems like Elastic, Jira, and cloud platforms
Build and operationalize AI agents for SRE workflows (incident triage, RCA, alert summarization, runbooks)
Drive AIOps maturity: alert correlation, anomaly detection, assisted RCA
Develop predictive models for capacity, failures, and incidents
Day-to-day Operations
On-call Production Support:
Provide production support on a shift according to the team on-call roster.
While not on call for production support, work on SRE projects and Tech support escalated and internal engineering/implementation team raised tickets
Work on SREs backlog items.
Leverage AI-assisted triage tools & MCP frameworks to prioritize alerts, detect anomaly patterns, and reduce noise during on-call
Continuously improve AI-driven incident routing and recommendation systems to optimize response efficiency.
Monitor and Maintain Systems:
Proactively monitor the health and performance of our services, systems, and infrastructure. Respond to alerts and incidents promptly to ensure high availability.
Effectively identifies & addresses monitoring and observability gaps
Implements effective alerting & notifications, minimizing false alerts
Creates and manages effective SRE Dashboards to report Key business metrics, SLAs, SLOs, SLIs & error budgets
Ensure SREs are meeting or improving on established SLOs
Proactively & effectively evaluates capacity planning to handle growth - scalability & traffic load
Contributes to innovative solutions like AI Assistant for proactive issue detection & response
Design and implement AI/ML-based anomaly detection for proactive identification of system degradation
Utilize predictive analytics models for capacity forecasting, incident prediction, and failure prevention.
Integrate AIOps platforms (like Elastic AI assistant) for intelligent alert correlation and root cause suggestions.
Develop self-learning monitoring systems that evolve with application behavior.
System reliability Improvements:
Actively participates and tracks execution of SRE projects aimed at improving system reliability
Effectively collaborates with cross teams to prevent reliability issues
Reviews change management tickets to identify and mitigate potential risks to system reliability
Ensure active participation in change activities and verify that accurate validations are performed by SRE & Engineering teams post implementation.
Participate in architecture reviews & assess the impact of architectural decisions on system reliability
Initiatives to perform chaos experiments to continuously learn and improve performance & stability of our systems
Contributes to innovative solutions that enhance system reliability & scalability
Drive initiatives to build self-healing systems using ML-based decision engines
Use AI models to simulate failure scenarios and predict system behavior under stress (intelligent chaos engineering)
Identify reliability risks using pattern detection across logs, metrics, and traces.
Contribute to building adaptive scaling systems using ML-based workload prediction.
Incident Management:
Actively participate in troubleshooting and resolving incidents, performing root cause analysis, Incident postmortems and implementing long-term fixes to prevent recurrence.
Acknowledge & quick recovery from incidents
Maintains quality of Root cause analysis (RCA) and corrective action plans
Proactively monitors, measures & adheres to optimal MTTR & MTTA requirements
Improves quality of SOPs, Adapts AI tools to reduce MTTR
Leverage AI-driven root cause analysis tools to accelerate incident diagnosis.
Implement automated incident summarization and timeline reconstruction using…

Skills asked for

Similar jobs

Apply on the employer's site

Your next role is already in here.

Search live openings from thousands of employers, save the ones worth a second look, and let JobBob keep watch for the rest.