woixi

Senior Dev Ops Engineer

 DevOps/SRE engineer with 5+ years of hands-on experience designing, automating and supporting cloud, on-premise and hybrid infrastructure. Expert in Kubernetes, CI/CD (GitLab CI, Jenkins, ArgoCD), IaC (Terraform, Ansible), monitoring (Prometheus, Grafana) and service security (SAST, SCA, Vault, WireGuard). Comfortable across the whole SDLC: I design, implement, optimize and document processes. Proven cases of accelerating releases and reducing MTTR by 30–50%. I work both with large enterprise solutions and with small businesses, and integrate easily into a team. Open to relocation and remote work, always focused on results and long-term quality. Contacts: email — [email protected], Telegram — @woixi. 


Experience: 5 years

Yearly salary: $50,000

Hourly rate: $40

Nationality: 🇹🇭 Thailand

Residency: 🇰🇿 Kazakhstan


Experience

System Engineer
Kaspersky
2026 - 2026
Maintained company-wide HashiCorp Vault and JFrog Artifactory in production: multi-datacenter HA clusters, 10M artifacts, used by every development team in the company. Area of responsibility: availability, incidents, operations automation. HashiCorp Vault, GitLab CI, Ansible Automated deployment of the Vault/Consul cluster (6 nodes, 3 datacenters, air-gapped segment): IaC pipeline in GitLab CI with validate, dry-run and deploy stages, targeting by node, role and datacenter. Removed static tokens from CI: a JWT (OIDC) from GitLab is exchanged for a short-lived Vault token (bound_claims, short TTL); KV v2 secrets are available only for the duration of the job. Automated engineers' access to nodes via Vault: users, SSH keys and sudo rules are created from secrets with a single vault kv put/delete command. Resolved Vault incidents (JWT/OIDC, AppRole, Kerberos, LDAP, leases, ACL policies) and automated the creation of policies, namespaces and roles from user requests. Linux baseline, CIS hardening Maintained an Ansible baseline for RHEL, Rocky Linux, CentOS, Debian and Ubuntu, accounting for network segments (HQ, DMZ, air-gapped perimeter) and repository mirrors in Artifactory. Adapted hardening to CIS Benchmarks (dev-sec/ansible-collection-hardening): around 60 parameters, including sysctl, auditd, PAM, SUID/SGID, modprobe, mount options and SELinux. Drove the air-gapped-segment rollout to its first successful run, fixing and documenting 11 incidents: KESL and KLNAgent updates, AppArmor and rsyslog on Ubuntu 24.04, duplicate imjournal logs, get_url hanging on the proxy. Deployed KESL and KLNAgent (Kaspersky Security Center), rsyslog with TLS forwarding, chrony and sshd. JFrog Artifactory, Python Resolved incidents across 10M artifacts: corrupted RPM and npm indexes, checksum mismatches in remote repositories (EINTEGRITY), stale metadata caused by the Nginx reverse-proxy cache. Wrote Python tooling on top of the REST API and AQL for artifact search, checksum verification and bulk operations. Stack: HashiCorp Vault, Consul, JFrog Artifactory, Ansible, GitLab CI, RHEL, Rocky Linux, CentOS, Debian, Ubuntu, SELinux, AppArmor, auditd, Python, Nginx, MSSQL, Kaspersky Security Center, rsyslog, Prometheus, Grafana
DevOps
Inline Group
2025 - 2026
Combined DevOps support of on-premise infrastructure with building an MLOps platform for a production computer vision system. Kubernetes, OKD, Deckhouse, Cilium Deployed and administered Kubernetes on bare metal: OKD 4 (CRI-O, OVN-Kubernetes, Routes, OLM, SCC) and Deckhouse (dhctl, NodeGroup, ingress-nginx, cert-manager, user-authn, cni-cilium modules). Configured HPA, rolling updates, probes, PodDisruptionBudgets and requests/limits, plus network isolation of environments using CiliumNetworkPolicy and Hubble. Piloted Talos Linux as an immutable OS for nodes: machine config, management through talosctl without SSH. Terraform, Ansible, CI/CD, DevSecOps Automated environment provisioning with Terraform and Ansible: setup time dropped from 1–3 days to 30–120 minutes. Ran CI/CD in Azure DevOps Server, Jenkins and GitLab CI; developed Helm charts with per-environment values. Embedded SAST (Static Application Security Testing) and SCA (Software Composition Analysis) into pipelines: remediation of critical vulnerabilities accelerated from 10–14 days to 2–4 days. Took part in the migration to RED OS, rewrote Dockerfiles to meet security requirements (multi-stage, non-root), set up Nexus Repository for images and artifacts. Prometheus, Grafana, SLI/SLO, Uptrace Maintained 99.9–99.95% SLA for on-premise services. Introduced SLI/SLO-based monitoring and alerting (Service Level Indicators/Objectives: availability, latency, error rate): MTTR (Mean Time To Recovery) decreased 3–6×, user-facing incidents dropped by 40–60%. Introduced distributed tracing with Uptrace (OpenTelemetry, ClickHouse). Apache Kafka Maintained Kafka brokers (partitioning, replication, consumer groups, consumer lag) and their integration with microservices. Designed the event-driven architecture of the CV system: video processing and business logic scale independently. MLOps: Triton, MLflow, Airflow, Kueue, GPU Took a computer vision system for monitoring railcar trains and wagon unloading to production: RTSP streams, OpenCV, PyTorch. Ran CV workloads on GPU nodes via the NVIDIA GPU Operator (device plugin, DCGM Exporter, MIG Manager). Set up inference on Triton Inference Server (ONNX, dynamic batching, automatic version pickup from MinIO) and a model registry on MLflow (PostgreSQL, MinIO). Built a retraining loop: drift is tracked in Evidently, low-confidence predictions go to labeling, and a retrain DAG in Apache Airflow registers and promotes the new version in MLflow. Configured GPU scheduling in Kueue: quotas for two teams in a shared cohort, borrowing and preemption. Compared MIG, MPS and time-slicing. Deployed LLM inference on TGI with batching tuning, Prometheus metrics and an OpenAI-compatible API. Stack: Kubernetes, OKD, Deckhouse, Talos Linux, Cilium, Docker, Helm, NVIDIA GPU Operator, Kueue, Terraform, Ansible, GitLab CI, Jenkins, Azure DevOps, Nexus Repository, Apache Airflow, MLflow, Triton Inference Server, TGI, Evidently, PyTorch, OpenCV, Apache Kafka, PostgreSQL, Redis, MinIO, Prometheus, Grafana, Uptrace, OpenTelemetry, RED OS
DevOps
GGR
2024 - 2024
Migration of infrastructure from Yandex Cloud to Hetzner and building a Kubernetes-based platform. Migration and infrastructure Migrated infrastructure from Yandex Cloud to Hetzner with minimal downtime: target environment preparation, data transfer, DNS cutover after verification. Described the Hetzner infrastructure in Terraform (servers, networks, firewall); node configuration was done with Ansible. Kubernetes (Deckhouse) Deployed a Deckhouse cluster and configured the ingress-nginx, cert-manager (Let's Encrypt TLS certificates) and prometheus modules. Set up unified sign-in to the cluster and web interfaces via the user-authn module (Dex with a GitLab OIDC provider); granted permissions through user-authz based on GitLab groups. Developed Helm charts and manifests for services. CI/CD and GitOps Deployed GitLab Runner with the Kubernetes executor. Image builds run through Kaniko, without a Docker daemon or privileged containers. Set up GitOps delivery with ArgoCD: an Application per service, automatic sync with prune and self-heal. Releases sped up 3–5×. Deployed Harbor as the image registry: projects, robot accounts for CI, vulnerability scanning via built-in Trivy, image retention policies. Data, monitoring, access Deployed a fault-tolerant PostgreSQL cluster on Patroni: automatic failover, streaming replication, automated backups and monitoring. Deployed RabbitMQ for event-driven communication between services. Set up monitoring and alerting on Prometheus, VictoriaMetrics for long-term metric storage, and Grafana dashboards. Deployed Vaultwarden for team password storage. Stack: Yandex Cloud, Hetzner, Terraform, Ansible, Kubernetes, Deckhouse, Helm, Docker, GitLab CI, Kaniko, ArgoCD, Harbor, Trivy, PostgreSQL, Patroni, RabbitMQ, Prometheus, VictoriaMetrics, Grafana, Vaultwarden
DevOps
MegaMailer
2022 - 2024
Described infrastructure in Terraform with remote state and separation of dev, stage and prod: networks, virtual machines, managed Kubernetes and object storage. In AWS deployed VPC, EC2, EKS and S3 with IAM-based access; in GCP: GKE, Compute Engine, Cloud Storage and service accounts; in Yandex Cloud: Managed Service for Kubernetes and Object Storage. In Hetzner Cloud spun up virtual machines for small services (Docker, Nginx, Let's Encrypt); standard server configuration was written as Ansible roles. Kubernetes, Helm, ArgoCD Migrated services from Docker Compose to Kubernetes and packaged them as Helm charts with per-environment values. Installed ingress-nginx, cert-manager and kube-prometheus-stack, set up GitOps delivery with ArgoCD. GitLab CI, Jenkins Wrote GitLab CI and Jenkins pipelines: tests, image build, publishing to a registry (Amazon ECR, Google Artifact Registry, Yandex Container Registry), updating the tag in the GitOps repository. HashiCorp Vault, Prometheus, Zabbix Deployed HashiCorp Vault for application secrets: KV v2, Kubernetes auth method, Vault Agent Injector. Set up cluster monitoring with Prometheus and Grafana with alerts to Telegram, and VM monitoring in Zabbix. Wrote deployment documentation so teams could maintain the infrastructure themselves. Stack: AWS, Google Cloud Platform (GCP), Yandex Cloud, Hetzner, Terraform, Ansible, Kubernetes, Helm, Docker, GitLab CI, Jenkins, ArgoCD, HashiCorp Vault, Prometheus, Grafana, Zabbix, Nginx
DevSecOps
Engineering Centre for Security systems
2021 - 2022
Monitoring and investigation of information security incidents, administration of security tools, integration of security checks into CI/CD. SIEM and incident response Administered the Wazuh SIEM (manager, indexer, dashboard, agents on hosts): wrote custom decoders and correlation rules. Configured file integrity monitoring (FIM), active response and vulnerability detection in Wazuh. Average incident investigation time dropped from several days to several hours. Responded to security incidents within SLA, developed and maintained response playbooks. Studied current attack techniques using the MITRE ATT&CK matrix and wrote detection rules and countermeasures for them. Analyzed network traffic (Wireshark, tcpdump) and identified anomalies, conducted penetration testing. Security tools and infrastructure Administered information security tools: EDR, NGFW, IDS/IPS, WAF, antivirus, unauthorized-access protection systems. Configured policies and rules, updated signatures, triaged alerts. Managed infrastructure services: Active Directory (users, groups, group policies), DNS, DHCP. Administered MySQL and MSSQL. Set up infrastructure monitoring in Zabbix: agents, templates, triggers. DevSecOps and automation Embedded SAST and SCA into Jenkins pipelines: time to fix detected vulnerabilities was reduced 2–4×. Automated routine operations in Bash, Python and Ansible, reducing manual work and error rates. Worked with container infrastructure on Docker, Kubernetes and Helm, storing images in Nexus Repository. Stack: Wazuh, SIEM, MITRE ATT&CK, EDR, NGFW, IDS/IPS, WAF, Wireshark, tcpdump, Jenkins, SAST, SCA, Docker, Kubernetes, Helm, Nexus Repository, Ansible, Bash, Python, Zabbix, Active Directory, DNS, DHCP, MySQL, MSSQL

Skills

aws
docker
gambling
gcp
git
kubernetes
linux
devops
english
javanese