Site Reliability Engineer
Akakçe
About Us Founded in 2000, Akakçe is Turkey’s leading price comparison platform, helping over 40 million users find the best deals every month. Our commitment to cutting-edge technology, data-driven insights, and continuous innovation makes us a key player in the e-commerce ecosystem. About the Role At Akakçe, we operate web and mobile platforms used by millions of users. The infrastructure underneath them is genuinely hybrid: physical servers and virtualization we own and run ourselves, Kubernetes and cloud services alongside them, and Linux and Windows workloads in the same estate. Someone has to keep all of it available, fast, and recoverable. We are looking for a Site Reliability Engineer to own that. You will work in the System team, reporting to the Head of System, alongside Technical Support. Experience with our exact tooling is useful, but we care far more about how you reason about failure, what you do in the first ten minutes of an incident, and whether you make problems stop coming back. This is not a role where you watch dashboards and file tickets. You own environments, you change them, and you are accountable for what happens when they break. What You Will Do Own and operate the production, stage, test, and development environments. Manage infrastructure end to end, from physical servers, network, and storage through to virtual machines and cloud resources. Build and maintain infrastructure as code, and run deployments through GitOps practices. Own and improve CI/CD pipelines; automate build, test, deployment, rollback, and recovery. Own DNS, load balancer, firewall, and WAF configuration. Own database operations, including provisioning, performance tuning, backup, and restore. Monitor infrastructure, platform, and service metrics, and decide what deserves an alert. Respond to incidents, run the postmortem afterwards, and drive the changes that stop a repeat. Plan capacity, and organize and execute planned maintenance. Own disaster recovery and failover testing, and prove that restores actually work. Support developers when a pipeline, deployment, or environment issue blocks them. Document infrastructure changes — no undocumented change reaches production. What We Look For We hire for engineering ability and operational judgment, not for a list of tools. We would like to talk to you if: You can walk through a production incident you handled — what you did, in what order, and why that order. You can explain a change you made to improve reliability, and how you knew afterwards that it worked. You have written a postmortem, and can say what actually changed as a result of it. You can describe a rollback you performed, and what made it safe to do. You have replaced a manual runbook with automation, and can explain what you did not automate and why. You can explain why an alert was noisy and what you did about the underlying cause, rather than muting it. You can read someone else’s pipeline or infrastructure code and understand it before you change it. You can say what you would check first when a service is slow but not down. Current Technologies We Use Our estate is broad, and nobody arrives knowing all of it. Depth in some of these counts for more than passing familiarity with all of them: Linux (Debian) and networking fundamentals; Windows Server in parts of the estate Containers and orchestration — Kubernetes, Docker, or comparable Infrastructure as code and GitOps — Terraform, ArgoCD, Helm, or comparable CI/CD pipelines — Jenkins, GitHub Actions, or comparable Observability — metrics and centralized logging, such as Grafana and Elasticsearch/Kibana Relational databases in production, including backup, restore, and high availability Edge and network services — reverse proxy, CDN, DNS, and WAF. Nice to Have Experience running high-traffic consumer platforms. Experience with on-premise infrastructure — physical servers, virtualization, and data center operations. Experience with database high availability and disaster recovery you have actually tested. Scripting and automation in any language you reach for — Bash, Python, Go, PowerShell. Experience migrating workloads between platforms without an outage. Infrastructure or automation work you can walk us through and explain in depth. How We Interview Our process is built around fundamentals and problem solving rather than tool trivia. Expect a practical exercise and a technical conversation about a real failure scenario: what you would check, in what order, what you would change, and what you would do differently the second time. We will ask you to explain your reasoning, and we will follow up on your answers. "Your personal data gathered in your application to this job advertisement will be processed in automated systems to manage job application procedures by Akakçe Bilgi Teknolojileri Sanayi ve Ticaret A. Ş. (“Akakçe”) according to the Law No. 6698 on the Protection of Personal Data (“KVKK”) and related regulations."
This job was verified from LinkedIn Jobs Türkiye. Applications are completed on the original source.
Apply on the original listing ↗
Something wrong with this job?