| Job Position and Company | Posted | Location | Salary | Tags |
|---|---|---|---|---|
|
| ||||
| $90k - $112k | ||||
| $86k - $91k | ||||
| $90k - $100k | ||||
ISO 9001 Certified | 400+ students | Learn more | by Metana | ||
|
| ||||
|
| ||||
| $90k - $150k | ||||
| $90k - $104k | ||||
| $91k - $109k | ||||
|
| ||||
|
| ||||
|
| ||||
| $90k - $112k | ||||
| $90k - $112k | ||||
| $121k - $123k |
Senior Infrastructure Engineer — IDC / Bare-Metal Kubernetes
About the Role
We are building our own IDC infrastructure from the ground up. As a Senior Infrastructure Engineer, you will lead the architecture, buildout, and operations of our colocation data center — from rack layout to the on-prem Kubernetes platform. You build it, you own it. This is a true greenfield role with no legacy baggage and a high degree of architectural ownership.
Your core expertise will center on three pillars: building physical infrastructure from zero, running production-grade self-hosted Kubernetes, and hybrid-cloud interconnect. As the platform matures, you'll help drive the next phase of cross-cloud and scale-out buildout.
Responsibilities
-
Design rack layout, network topology, and overall configuration for colocation
-
Design and implement an isolated Out-of-Band (OOB) management network (BMC, IPMI, Redfish) with full security hardening
-
Coordinate with colocation and hardware vendors — servers, switches, cross-connects, uplinks, remote hands, and cage/rack administration
-
Build bare-metal automation for fleet-wide zero-touch OS provisioning (MAAS / Tinkerbell / Cluster API)
-
Set up production-grade self-hosted Kubernetes clusters from scratch and maintain the core stack: Cilium (CNI), GitOps (ArgoCD), etc.
-
Own cluster lifecycle: upgrades, scaling, node maintenance, incident response
-
Write SOPs, runbooks, and post-mortems; participate in 24×7 on-call rotation
-
Establish hybrid-cloud interconnect between IDC and public cloud (VPN / Direct Connect / peering)
-
Pave the way for scale-out: GPU workloads (NVIDIA GPU Operator, passthrough, MIG) and VM workloads running alongside containers (KubeVirt or similar)
Requirements
-
5+ years in data center / infrastructure / platform engineering
-
Hands-on experience building physical infrastructure from scratch: rack layout, network topology, server commissioning, and coordinating cross-connects and remote hands with colo / vendors
-
Practical experience designing and operating OOB management networks (BMC, IPMI, Redfish)
-
Have stood up production-grade self-hosted Kubernetes from scratch, and can independently debug cluster-level issues (CNI, CSI, storage)
-
Strong Linux systems administration and performance tuning (kernel, networking, storage I/O)
-
Bare-metal automation experience with at least one of: MAAS, Tinkerbell, Cluster API
-
Proficient with Terraform, Ansible, and at least one scripting language (Python / Go / Bash)
-
Experience with Cisco network and related techniques (VLAN, LACP/LAG, BGP, ACL, etc)
-
Experience with Palo Alto firewall configuration
-
Experience with storage systems (NetApp, Dell EMC, Pure Storage)
-
Fluent in English or Mandrain
Nice to Have
-
High-density racks (30kW+) and 400G+ networking experience
-
Familiarity with immutable OS (Talos Linux / Flatcar / Bottlerocket)
-
Proficiency across both AWS and GCP; cross-cloud data migration experience
-
Experience building storage or Bigdata / offline data clusters
-
Virtualization experience (KubeVirt or similars)
-
Exposure to NVIDIA GPU Operator and K8s GPU workloads
-
CKA / CKS certification
-
CCNP / CCIE certification
-
Professional-level Japanese