alexdevops

Senior Dev Ops Engineer

 DevOps/MLOps engineer with 7.5 years of experience building and maintaining high-load systems in e-commerce and fintech. Specialized in Kubernetes, MLOps, CI/CD, and observability. Built an ML platform from scratch at Wildberries, reducing model deployment time from days to 2–3 hours. Managed Kubernetes clusters (>1000 workers) at Alfa-Bank, ensuring 99.95% uptime for critical banking services. Expert in Kubernetes, Terraform, GitLab CI, Prometheus, and GPU infrastructure. Looking for a Senior DevOps/MLOps role where I can apply my skills to build reliable, scalable platforms.
Also has experience in crypto exchange infrastructure: managed Kubernetes clusters for high-load trading services, maintained blockchain nodes, and implemented DDoS protection 


Experience: 6 years

Yearly salary: $60,000

Hourly rate: $35

Nationality: 🇧🇾 Belarus

Residency: 🇧🇾 Belarus


Experience

ML Engineer
Wildberries
2024 - 2026
Designed and maintained a Kubernetes platform for ML model training, batch jobs, and online inference. Supported the full ML lifecycle: training, testing, registration, deployment, and monitoring. Deployed ML services via Helm and ArgoCD using GitOps. Configured HPA and KEDA for autoscaling of inference services based on technical and business metrics. Managed resource requests/limits, quotas, affinity, tolerations, and PodDisruptionBudgets. Operated NVIDIA GPU nodes, configured device plugins, and planned GPU workloads. Diagnosed PCI passthrough, NUMA, CUDA, driver, and GPU availability issues. Participated in capacity planning and optimization of CPU, RAM, and GPU usage. Developed CI/CD pipelines in GitLab CI for building, testing, and releasing ML services. Created universal Helm templates for standardized application deployment. Automated server and Kubernetes node configuration with Ansible. Supported Docker Registry and Nexus, managed image versions and lifecycle. Automated health checks and routine operational tasks using Bash. Developed monitoring with Prometheus, Grafana, Loki, and Tempo. Configured dashboards and alerts for Kubernetes, GPU nodes, ML pipelines, and inference services. Monitored latency, throughput, error rate, saturation, and availability. Centralized log collection and analysis for distributed applications. Participated in on-call, incident diagnosis, and RCA. Developed runbooks for service recovery and troubleshooting. Analyzed issues at Kubernetes, Linux, network, storage, and message broker levels. Used Kafka as event transport for online and batch data processing. Monitored consumer lag, broker health, and asynchronous operations. Worked with PostgreSQL, ClickHouse, Redis/KeyDB, and S3-compatible storage. Operated OpenStack cloud infrastructure. Diagnosed compute, network, and storage issues. Collaborated with Data Scientists, ML engineers, developers, and infrastructure teams.
DevOps Engineer
Alfa-Bank
2021 - 2024
Administered HighLoad Kubernetes clusters on OpenStack (Kubespray) and Yandex Cloud. Configured zero-downtime releases via ArgoCD using Canary and Blue-Green strategies, tuned Graceful Shutdown and PDBs. Configured KEDA for microservice autoscaling based on Kafka Consumer Lag during peak loads. Configured Calico CNI, microsegmentation via Network Policies for critical banking segments. Declared hybrid infrastructure via Terraform + Terragrunt (modular DRY architecture, strict code review). Configured HashiCorp Vault and Vault Agent Injector (sidecar pattern). Configured mTLS between services. Developed unified Helm charts. Integrated automated security checks (SAST, Trivy image scanning, non-root containers). Migrated from Prometheus to VictoriaMetrics Cluster for high-cardinality metrics. Built logging system based on Vector + OpenSearch, supported ELK Stack. Organized monitoring of business metrics and SLA in Grafana, conducted RCA (OOMKilled, CPU Throttling, I/O bottlenecks), wrote post-mortems. Operated PostgreSQL clusters (Patroni + HAProxy + etcd). Configured backups via WAL-G to S3, conducted DR drills. Supported high-load Apache Kafka clusters (KRaft migration), controlled partitioning and resolved event processing delays.
DevOps Engineer
Bitfinex
2019 - 2021
Managed Kubernetes clusters for high-load crypto exchange services (spot trading, futures, wallet). Deployed and maintained blockchain nodes (Bitcoin, Ethereum, BSC) and monitored their health. Built CI/CD pipelines in GitLab CI for trading engine, matching engine, and wallet services. Configured Prometheus and Grafana for 24/7 monitoring of exchange infrastructure and blockchain nodes. Implemented DDoS protection at L3/L4 and L7 levels for public API and web endpoints. Worked with Kafka, PostgreSQL, Redis, and S3-compatible storage for real-time transaction processing. Automated infrastructure provisioning with Terraform and Ansible. Collaborated with security team on cold/hot wallet infrastructure and key management. Participated in on-call rotation and incident response for critical trading outages.

Skills

ansible
ci-cd
devops
docker
gpu
grafana
kubernetes
postgres
python
redis
terraform
english
russian