…

AI/DSP Kernel Optimization Engineer

MulticoreWare Pvt Ltd · IT Services & Consulting

  • Chennai, India
  • On-site
  • Posted 20 days ago
  • Software Engineering
  • Full time

About the job

We are looking

for an AI/DSP Kernel Optimization Engineer to develop and optimize AI inference

kernels, analyze application performance, and integrate optimized kernels into

AI inference pipelines. The role involves low-level C/C++ optimization,

SIMD/vectorization, profiling, debugging, and close collaboration with hardware

and software teams to improve overall inference performance.

Responsibilities:

Develop and optimize AI inference kernels using

C/C++.

Optimize kernels for different data types,

including INT8, INT16, BF16, and FP32.

Perform SIMD/vectorization, intrinsics, memory

optimization, and performance tuning.

Profile and analyze application performance to

identify optimization opportunities.

Debug and resolve functional and performance

issues.

Integrate optimized kernels into the AI inference

pipeline.

Work closely with hardware and software teams to

improve overall inference performance.

Requirements

Education:

B.E.,

B.Tech., M.E., M.Tech., or equivalent.

Technical Skills (Must haves):

Good knowledge of C/C++ programming.
Basic understanding of AI/Deep Learning models

and inference.

Experience with DSPs, hardware accelerators, or

similar compute platforms.

Knowledge of SIMD/vectorization, intrinsics, and

low-level software optimization.

Good profiling, debugging, and performance

analysis skills.

Need to have (Can be bridged):

Understanding of computer architecture and memory

optimization.

Good to have (Not essential):

Experience in AI/ML inference optimization.
Knowledge of different numerical data types such

as INT8, INT16, BF16, and FP32.

Familiarity with performance profiling and

benchmarking tools.

Preferred Qualifications (Optional):

Experience with C7x DSP, HWA-MMA, or similar

hardware accelerators.

Hands-on experience optimizing inference

workloads for DSPs or specialized compute platforms.