Position: AI Infrastructure System Performance Architect (IRT815ST RM 4407)
Position Summary
Analyze system-level performance, scalability, and bottlenecks of the company’s chiplet-based AI infrastructure using virtual platforms and performance models.
Education : B.E. / B.Tech / M.E. / M.Tech in Electronics, Electrical, Computer Engineering, or related discipline.
Key Responsibilities
- Memory / Data-Movement Modelling: Analyze memory bandwidth and latency requirements, model XPU-to-memory and chiplet-to-memory traffic, and identify data-movement bottlenecks.
- Multi-Chiplet / Multi-XPU Scalability: Analyze scaling from single-XPU to multi-XPU systems, evaluate scale-up/scale-out architectures, and study bandwidth, latency, congestion, and resource utilization.
- AI Infrastructure System Modelling: Develop system-level performance models covering compute, interconnect, memory, and I/O interactions, and develop representative AI/HPC traffic patterns.
- System Performance & Architecture Exploration: Perform architecture trade-offs and analyze latency, bandwidth, throughput, utilization, queue depth, congestion, and bottlenecks; generate performance reports and architecture recommendations.
Required Technical Skills
- 8–15+ years of system/performance architecture experience.
- AI/HPC infrastructure and performance modelling.
- SystemC/TLM.
- C++ / Python.
- PCIe / CXL / UCIe.
- HBM/DDR.
- System-level architecture.
Preferred Skills
- XPU/GPU/NPU architecture.
- AI workload characterization.
- UALink and networking.
- Cluster/rack-scale architecture.
- Power-performance analysis.
Project Experience : Demonstrated system-level performance analysis (latency, bandwidth, throughput, utilization, congestion) on a chiplet-based or multi-XPU AI system.
Soft Skills
- Strong analytical rigor in translating raw performance data into architecture recommendations.
- Ability to collaborate across connectivity, memory, and power architecture teams to build a coherent
system view.
Good to Have
- Experience characterizing real AI/HPC workloads for use in performance models.
- Exposure to cluster or rack-scale system architecture.
Expected Deliverables
- System-level performance models covering compute, interconnect, memory, and I/O.
- Architecture trade-off studies and performance reports with recommendations.
- Representative AI/HPC traffic pattern definitions for use across the team’s models.
Growth Path
- Progress into Principal Performance Architect, System Architecture Lead, or Practice Lead (Performance Engineering) roles.
*****************************************************************************************************************
Apply for this position
Mention correct information below. Mention skills aligned with the job description you are applying for. This would help us process your application seamlessly.
