Staff Software Engineer - Reporting, Data Platform & Observability
Onetrust · Atlanta, Georgia · 2026-09-17
About this role
Strength in Trust
OneTrust’s mission is to enable innovation through the responsible use of data and AI. We believe that ensuring data is trusted shouldn’t slow teams down—it should accelerate what’s possible. This led us to develop the first technology platform for responsible data use in 2016. Today, with AI representing the latest and most impactful expansion of data yet, OneTrust is once again redefining what responsible innovation looks like. OneTrust, the AI‑Ready Governance Platform™, unifies regulatory intelligence, automation, and connected governance workflows so businesses can continue to move at the speed of AI while ensuring good governance to prevent data misuse at scale. Trusted by thousands of organizations worldwide, OneTrust is shaping the future where trusted data becomes a transformative force for business and society.
The Challenge
OneTrust is seeking a Staff Software Engineer to join the Reporting and Data Platform team. This is a hands-on individual contributor role focused on designing, building, operating, and improving distributed backend services and data-processing platforms.
You will work across Java microservices and Python/PySpark data pipelines, with a strong focus on reliability, scalability, performance, and observability. You will take complex or ambiguous problems from investigation through production delivery and help improve the systems that power reporting and data-driven experiences.
Your Mission
Technical Ownership and Delivery
• Own complex features and technical improvements from discovery through production rollout, making substantial hands-on contributions across backend services and data-processing pipelines.
• Investigate ambiguous problems, identify root causes, evaluate trade-offs, and implement pragmatic solutions that improve code quality, maintainability, automated testing, and operational readiness.
• Review code and technical designs, document important implementation decisions and system behavior, and partner with product managers, engineers, and other teams to clarify requirements and deliver outcomes.
• Apply AI-assisted engineering tools such as Devin, Claude, or similar systems to accelerate delivery while maintaining production-quality design, code, tests, security, and operational readiness.
Backend and Distributed Systems
• Design and implement production services using Java, Spring Boot, and Maven, including APIs, asynchronous workflows, report generation, aggregation, export, and scheduling capabilities.
• Develop event-driven functionality using Kafka and related messaging patterns, and work with caching technologies, relational storage, and service-to-service integrations.
• Improve service performance, scalability, fault tolerance, and resource efficiency through appropriate patterns for retries, idempotency, caching, backpressure, concurrency, and failure recovery.
• Diagnose issues across services, queues, databases, and downstream dependencies, and modernize established capabilities incrementally while maintaining production stability.
Data Engineering
• Build and maintain ingestion and transformation pipelines using Python, PySpark, Azure Databricks, and Delta Lake across batch and streaming workloads.
• Implement schema evolution, checkpoint management, deduplication, replay, late-arriving-data handling, and standardized data-layer patterns.
• Optimize Spark joins, partitioning, Delta operations, cluster utilization, and query performance while troubleshooting failed, delayed, or inefficient Databricks workloads.
• Protect tenant boundaries across joins, aggregations, deduplication, and Delta operations; implement data-quality controls; and monitor data freshness, completeness, and correctness.
• Work securely with Azure storage, identities, secrets, and encryption mechanisms.
Observability, On-Call, and Operational Excellence
• Improve observability across backend services, event-driven workflows, and data pipelines using meaningful metrics, structured logs, traces, and business telemetry.
• Build and maintain actionable dashboards, monitors, and alerts using Datadog and Grafana, applying OpenTelemetry, Prometheus, and Micrometer patterns where appropriate.
• Participate in the on-call rotation and incident-response workflows, using PagerDuty, Datadog monitors, or equivalent platforms to diagnose production issues and drive sustainable resolution.
• Reduce recurring alerts and operational toil by improving alert quality, eliminating noisy or non-actionable monitors, creating runbooks and diagnostic tools, and implementing corrective actions from blameless incident reviews.
• Improve end-to-end correlation and monitor availability, error rates, latency, ingestion lag, data freshness, event throughput, consumer lag, job health, rejected records, checkpoint health, tenant-specific failures, data-quality violations, and Spark resource utilization.
What Success Looks Like
• You require limited direction after understanding the desired outcome and relevant constraints, and you break ambiguous problems into concrete, deliverable work.
• You own work through design, implementation, testing, deployment, production validation, and ongoing operation.
• You use production evidence and telemetry to prioritize improvements and resolve root causes rather than repeatedly treating symptoms.
• You reduce alert volume and operational toil over time without hiding genuine system risks, leaving systems easier to operate after each incident.
• You make sound trade-offs among delivery speed, reliability, performance, security, cost, and maintainability while collaborating constructively without formal authority.
You Are
You are a self-directed, hands-on Staff Engineer who enjoys solving complex problems across distributed services and data platforms. You think in systems and trade-offs, take ownership of production behavior, and use clear design thinking to simplify solutions and reduce code-delivery…
Skills asked for
- java
- microservices
- python
- spring boot
- kafka
- azure
- databricks
- spark
Similar jobs
- Staff Software Engineer - Consent & PreferencesOnetrust · Atlanta
- Staff Software Engineer - AI GovernanceOnetrust · Atlanta
- Staff Software Engineer - Developer ExperienceOnetrust · Atlanta
- Staff Software Engineer - AI Governance IntegrationsOnetrust · Atlanta
- Staff Wireless Software EngineerAxon · Atlanta
- Staff Software EngineerAxon · Atlanta
- Staff Software EngineerAndurilindustries · Atlanta
Your next role is already in here.
Search live openings from thousands of employers, save the ones worth a second look, and let JobBob keep watch for the rest.