Avance Consulting
Site Reliability Engineer
Die Stellenanzeige deiner ursprünglichen Anfrage ist
leider nicht mehr vakant.
Wie wäre es mit diesem
Job-Angebot?
Position Overview
- Security Clearance & Vetting Level: Public Sector Clearance + NdK (Nachweis der Kundigkeit)
The Observability & SRE Engineer builds and operates central telemetry stacks to provide visibility across distributed cloud infrastructure. This role implements metric collection, log aggregation, distributed tracing, and standalone alerting tailored for Air-Gap operations.
Key Responsibilities
- Monitoring Stack Setup: Deploy and maintain Prometheus, Grafana, Loki, and Jaeger stacks as code.
- Dashboards & Alerting: Develop PromQL/LogQL monitoring dashboards; configure AlertManager routing and inhibition for autonomous Air-Gap operations.
- Distributed Tracing: Implement OpenTelemetry Collectors and SDKs for auto-instrumentation and OTLP trace propagation.
- SIEM Integration: Configure secure log export and security event management (CEF/Syslog) for external SIEM platforms.
Technical Qualifications & Skills
- Must Have:
- Deep expertise in Prometheus (PromQL, ServiceMonitor, Federation, Remote Write) and Grafana.
- Hands-on experience with Loki log aggregation and AlertManager routing.
- Proficiency with OpenTelemetry (Collectors, SDKs, OTLP) and Jaeger distributed tracing.
- Knowledge of SIEM integrations and security event logging.
- Scripting skills in Go, Python, Shell, and YAML.