JobBobsReal-time global job discoveryLive

Senior Site Reliability Engineer

Order.co · United States · 2026-10-07

seniorRemote
Apply on the employer's site

About this role

Order.co is the System of Action for the Office of the CFO, transforming the way businesses purchase and pay into an intuitive, B2C-like shopping experience. Order.co leverages embedded AI agents and embedded financial products to reinvent the way businesses connect with their vendors.

End users enjoy a seamless, zero-training buying experience, while finance and procurement leaders gain a single platform to orchestrate how the business “should operate”. The result is an all-in-one solution that serves as a gravitational pull for spend and data, automating and eliminating procurement and finance workflows from requisition to reconciliation along the way.

Order.co is on the cutting edge of B2B Agentic Commerce, poised to be the market leader in creating a more predictive, prescriptive, and personalized experience for users.

Founded in 2016 and headquartered in New York City, Order.co oversees nearly half a billion in annualized spend across hundreds of customers like WeWork, SoulCycle, Lume, and [solidcore]. Order.co has raised $75M in funding from industry-leading investors like MIT, Stage 2 Capital, Rally Ventures, 645 Ventures, and more. Order.co has been proudly named a 50 to Watch by Spend Matters and a Best Place to Work by BuiltIn and Inc. Magazine.

The Role

As a Senior Site Reliability Engineer on the Platform team, you will ensure that software systems are reliable, scalable, performant, and operationally efficient. You blend software engineering skills with infrastructure and operations expertise to keep critical systems running smoothly while enabling rapid product development.

Responsibilities

Reliability Engineering & Infrastructure Ownership

• Design, build, and operate highly available, scalable, and fault-tolerant infrastructure and platform services

• Own reliability, availability, latency, and operational excellence for critical production systems and services

• Define and maintain service level objectives (SLOs), service level indicators (SLIs), and error budgets across platform systems

• Lead incident response efforts for complex production outages; drive root-cause analysis and long-term remediation actions

• Build resilient systems that gracefully handle failures, traffic spikes, dependency degradation, and regional outages

• Continuously improve system reliability through automation, observability, performance tuning, and capacity planning

Automation & Platform Engineering

• Develop infrastructure automation and self-service tooling to reduce operational toil and improve engineering velocity

• Build and maintain CI/CD pipelines, deployment automation, and release engineering workflows

• Implement infrastructure as code (IaC) practices using tools such as Terraform, CloudFormation, and container orchestration

• Improve developer experience by building reliable internal platforms, operational tooling, and standardized deployment patterns

• Drive adoption of GitOps, immutable infrastructure, and automated remediation patterns

Observability & Operational Excellence

• Design and maintain comprehensive monitoring, logging, tracing, and alerting systems for distributed services

• Establish actionable alerting standards that reduce noise while improving incident detection and response times

• Analyze production trends, system bottlenecks, and failure patterns to proactively prevent incidents

• Lead operational readiness reviews, disaster recovery planning, and game-day exercises

• Improve mean time to detect (MTTD) and mean time to recovery (MTTR) through tooling, automation, and process refinement

Systems Architecture & Scalability

• Participate actively in architecture and infrastructure design reviews

• Propose scalable and reliable platform designs that account for multi-region deployment, redundancy, failover, and security considerations

• Evaluate trade-offs between reliability, scalability, operational complexity, and engineering velocity

• Identify systemic risks and operational gaps before they become production incidents

• Partner with engineering teams to ensure services are designed with operability, observability, and resilience in mind from day one

Security & Compliance

• Approach infrastructure and operational practices with a strong security mindset

• Implement and maintain secure cloud networking, secrets management, IAM policies, and infrastructure hardening standards

• Partner with Security and Compliance teams to ensure systems meet organizational and regulatory requirements

• Drive operational best practices around vulnerability management, patching, and production access controls

End-to-End Ownership & Collaboration

• Scope and estimate infrastructure and reliability initiatives accurately

• Coordinate production rollouts, maintenance events, and reliability improvements across teams

• Communicate operational risks, dependencies, and incident impacts clearly to technical and non-technical stakeholders

• Collaborate closely with Software Engineering, Security, Product, and Operations teams to improve platform reliability and scalability

• Serve as a trusted escalation point during critical production incidents

Mentorship & Technical Leadership

• Mentor junior and mid-level engineers on reliability engineering principles, operational excellence, and infrastructure best practices

• Raise the operational maturity of the engineering organization through documentation, reviews, and technical guidance

• Drive improvements in team standards around observability, incident management, automation, and infrastructure design

• Influence technical decisions through credibility, operational expertise, and strong engineering judgment

Qualifications

• You are motivated by accountability — you own outcomes, not just tasks

• You are results-oriented and measure success by shipped, working software

• You are motivated by correctness in code that touches money — the consequences of a bug land on real customer balances,…

Skills asked for

Similar jobs

Apply on the employer's site

Your next role is already in here.

Search live openings from thousands of employers, save the ones worth a second look, and let JobBob keep watch for the rest.