IT is the engineering that builds and runs an organization's technology estate: data centers and AI data centers, IT Infrastructure, Enterprise IT platforms, Cloud Infrastructure, enterprise networks, and the production operations that keep all of it available. The work spans systems administration, virtualization, storage, ERP and CRM systems, IT service management, containers and Kubernetes, infrastructure as code, observability, incident response, and high-performance computing. The market keeps expanding. IDC measured USD 318 billion of worldwide AI infrastructure spending in 2025, more than double the 2024 total, and forecasts USD 497 billion for 2026 . The FinOps Foundation's 2026 survey of practices managing over USD 83 billion in cloud spend found that 98% now manage AI spend, up from 31% two years earlier .
Challenges in IT Recruiting
Cloud optimization and repatriation cycles rebalance role demand
Cloud cost discipline remains the defining operating pressure on infrastructure teams. Flexera's 2025 survey of 759 cloud decision-makers found 21% of cloud workloads repatriated in the past year while more than half remained in public cloud with further migration expected, and 59% of organizations now run a dedicated FinOps team, up from 51% . FinOps Foundation's 2026 data shows the discipline spreading beyond public cloud: 90% of practices manage SaaS or plan to, 64% manage licensing, 57% manage private cloud and 48% manage data center spend. Optimization is yielding diminishing returns, with practitioners reporting they have hit the big rocks of waste and now face smaller, harder opportunities . Migration-era cloud architects give way to engineers who can place workloads across public cloud, private cloud and IT Infrastructure, read unit economics, and run the commitments that govern spend before it is committed.
AI data center buildout pulls infrastructure and HPC talent
Density is the operational story. IDC put AI infrastructure spending at USD 318 billion in 2025, more than double 2024, and raised its 2026 forecast to USD 497 billion after a first quarter that grew 33.1% year over year . That capital arrives as physical constraint: Uptime Institute's 2026 survey of more than 800 data center owners and operators found a growing share reporting peak rack densities of 30 kW or higher . The hiring consequence is scarcity at the joins. GPU cluster architecture, High-Performance Computing scheduling, InfiniBand interconnects and parallel file systems are thin markets, as are Network Engineering for east-west traffic and capacity planners who model power against workload. Kubernetes sits underneath much of it: 82% of container users run it in production, and it has become the common layer for AI workloads . An AI data center needs integrated power, network, storage and platform engineering.
Platform-as-a-product and error budgets rewrite operations titles
Platform engineering has become the organizing idea for how operations teams deliver. DORA's 2025 study, drawn from nearly 5,000 technology professionals, found 90% of organizations have adopted at least one platform and that AI amplifies the system it lands in, lifting throughput while worsening stability where foundations are weak . CNCF's annual survey adds that Kubernetes is production-standard for 82% of container users, and its emphasis has moved from technical complexity to people, culture and organizational alignment . Title inflation is the predictable side effect. A DevOps engineer may maintain CI pipelines for one team or build self-service infrastructure for hundreds of developers. A Site Reliability Engineering (SRE) title may mean on-call governed by service-level objectives and error budgets, or operations under a fashionable name. Establish whether the seat owns delivery, reliability or platform as a product before writing the specification.
Legacy modernization keeps mainframe and enterprise integration skills scarce
The mainframe is not disappearing; the shift is generational. BMC's 20th annual survey of more than 1,000 mainframe practitioners found 66% now identify as millennials or Gen Z, up from 37% in 2018, 65% already use generative AI with the platform, and AIOps has become a top-three operations priority . Modernization rather than replacement demands a broader skill mix than the platform's reputation suggests. Kyndryl's research found 89% rate the mainframe as extremely or very important to strategy, with 56% of mission-critical applications still running there, yet 45% say a lack of cybersecurity skills hurts modernization and 41% lack platform-specific capabilities such as z/OS Connect and Zowe . Enterprise IT carries a parallel load as ERP and CRM platforms move to cloud or hybrid deployments, lifting demand for integration architects fluent in both the legacy and the modern service model. The scarce profile is not a COBOL relic hunter but an engineer fluent in both worlds.
Cyber-resilience regulation makes operations evidence a hiring criterion
Regulation has turned operational resilience into a legal obligation. DORA has applied since 17 January 2025 to 20 types of financial entities and their ICT third-party providers, covering risk management, resilience testing and incident reporting . NIS2, with a transposition deadline of 17 October 2024, brings cloud computing, data center and managed service providers into scope, requires staged incident reporting at 24 hours, 72 hours and one month, and allows fines of at least EUR 10 million or 2% of worldwide turnover for essential entities . Uptime's 2025 analysis found IT and networking issues caused 23% of impactful outages, procedure-related human error rose ten percentage points year over year, nearly 40% of organizations suffered a major human-error outage in three years with 85% traceable to procedure failures, and third parties account for about two-thirds of publicly reported outages over nine years . Uptime's 2026 survey adds that one in ten outages remains serious or severe and costs continue to climb . Employers now need reliability evidence, not reliability vocabulary.
Remote teams and pay transparency reshape the offer
Operations teams are no longer bounded by a single site. Coverage follows the sun, on-call rotations are distributed, and candidates compare roles across national markets, making published pay a first-order recruiting variable. Under the EU Pay Transparency Directive, employers must inform job seekers of the starting salary or pay range in the vacancy notice or ahead of interview, may no longer ask about pay history, and organizations with at least 100 employees must publish gender pay gap information . Competition for operators remains severe: Uptime's 2026 survey found more than half of operators struggle to find qualified candidates, and turnover persists as engineers are hired away by other data center companies . Transparency shifts the contest onto what actually retains infrastructure engineers: the scale and modernity of the estate, the quality of the on-call experience, incident load, and whether the platform roadmap is credible. A vague advert with a hidden range no longer competes.
SRE, DevOps and platform titles hide different estates
Shared vocabulary is the defining assessment problem here. Kubernetes, observability, automation and reliability appear on nearly every operations CV, yet the work behind them differs by estate, scale and stack. A systems administrator patching a virtualized server room and a platform engineer running multi-tenant clusters across three clouds both claim infrastructure as code; one owns a few hundred virtual machines, the other owns blast radius. Below the surface: managed Kubernetes distributions against self-managed control planes, Terraform and Crossplane against Ansible-first estates, open-source pipelines against commercial suites, on-premises arrays against parallel file systems feeding GPU training. CNCF's data shows how standard Kubernetes has become, which is why the keyword discriminates so poorly .
The test is specificity. Ask which estate the candidate operated, how many clusters fell under their change control, which distribution and IaC toolchain they owned, what their largest incident was, and which service-level objectives their platform reported against. Weak answers stop at the tool name; strong answers name architectures, numbers and consequences.
Incidents, SLOs and estates a vendor list cannot prove
Candidate claims are easy to inflate and expensive to falsify late. A CV can list service-level objectives without a negotiated error budget, claim FinOps fluency without renegotiating a commitment, or describe an AI data center build without touching power limits. Verification means reconstructing ownership: what the candidate decided, what changed in production, and what the baseline was before and after. FinOps Foundation's data shows why cost skills cannot be guessed at: AI cost management is the capability teams most want and the hardest to assess when provider pricing varies and value is hard to attribute . A senior operations mis-hire is discovered during an incident, when the cost of correction is highest and the blast radius is real.
Assessment in this sector therefore has to be technical and evidence-based. The questions that separate candidates are concrete: whether a reliability engineer has carried production on-call under published objectives, whether a network engineer has designed a fabric and debugged it, whether a platform engineer has migrated clusters without customer-visible downtime, whether an HPC engineer has improved throughput on a constrained fabric, and whether a FinOps-aware architect has changed a budget outcome. A process that cannot test those claims forwards fluent CVs to panels whose time is the scarcest resource. Classification comes before sourcing: fix the estate, the operating model and the evidence that proves capability, then search the platform teams, provider ecosystems, enterprise IT shops and research computing groups where that evidence lives. Assessment quality decides whether the seat is filled by someone who can run infrastructure the business depends on.
References
- AI Infrastructure Spending Holds Near $90 Billion in Q1 2026 as ARM Overtakes x86 in Accelerated Servers; 2026 Forecast Raised to $497 Billion — IDC. (accessed 2026-09-18)
- The State of FinOps 2026 — FinOps Foundation. (accessed 2026-09-18)
- The latest cloud computing trends: Flexera 2025 State of the Cloud Report — Flexera. (accessed 2026-09-18)
- 16th Annual 2026 Global Data Center Survey: Deployment of High Density Racks Rising Fast, Operators Face Continued Recruiting and Retention Pressures — Uptime Institute. (accessed 2026-09-18)
- The CNCF Annual Cloud Native Survey: The Infrastructure of AI's Future — Cloud Native Computing Foundation (CNCF). (accessed 2026-09-18)
- Announcing the 2025 DORA Report: State of AI-Assisted Software Development — DORA, Google Cloud. (accessed 2026-09-18)
- 20th Annual BMC Mainframe Survey Reveals Confidence, Generational Shift in Views on Mainframe Development — BMC Software. (accessed 2026-09-18)
- The mainframe skills gap: What organizations often miss — Kyndryl. (accessed 2026-09-18)
- Digital Operational Resilience Act (DORA) — European Insurance and Occupational Pensions Authority (EIOPA). (accessed 2026-09-18)
- Directive on measures for a high common level of cybersecurity across the Union (NIS2 Directive) - FAQs — European Commission. (accessed 2026-09-18)
- Uptime Announces Annual Outage Analysis Report 2025 — Uptime Institute. (accessed 2026-09-18)
- New EU rules on pay transparency explained — European Commission. (accessed 2026-09-18)
