…

SSE - Optimization Engineer

MulticoreWare Pvt Ltd · IT Services & Consulting

  • Ramapuram, India
  • On-site
  • Posted about a month ago
  • Software Engineering
  • Full time

About the job

We are looking

for a Senior Software Engineer to develop and optimize deep learning models,

including CNNs, LLMs, and MoE, for efficient inference across CPU, GPU,

hardware accelerators, and edge devices. The role focuses on quantization,

model compression, high-performance kernel implementation, transformer

optimization, and production deployment.

Responsibilities:

Develop and optimize deep learning models (CNNs,

LLMs, MoE) for efficient inference across CPU, GPU, hardware accelerators, and

edge devices.

Design and implement quantization algorithms

(PTQ, QAT, GPTQ, AWQ) from scratch.

Apply model compression techniques such as

pruning, decomposition, and distillation.

Implement and optimize quantized kernels (INT8,

INT4, FP8) using C++ for high performance.

Translate research papers into production-ready

implementations.

Optimize latency, throughput, and memory usage

for real-world deployment.

Work on transformer optimization including

KV-cache, PEFT (LoRA/QLoRA), and MoE models.

Profile, benchmark, and debug model performance

across different hardware platforms.

Collaborate with ML, compiler, and hardware

teams to deliver optimized solutions.

Requirements

Education:

BE/BTech/MS/MTech in Computer Science or a related field.

Technical Skills (Must haves):

4+ years of relevant experience.
Strong programming skills in Python and C++.
Proven experience in quantization algorithms

(PTQ, QAT, GPTQ, AWQ).

Hands-on experience in pruning, model

compression, and inference optimization.

Experience implementing quantization or

optimization techniques from scratch.

Strong understanding of CNNs, Transformers, and

LLM architectures.

Experience with PyTorch / ONNX and model

deployment pipelines.

Strong problem-solving and performance

optimization skills.

Need to have (Can be bridged):

No

additional bridged skills were specified.

Good to have (Not essential):

Experience with MoE architectures and PEFT

techniques (LoRA, QLoRA).

Knowledge of TensorRT, ONNX Runtime, TVM, and

MLIR.

Familiarity with hardware-aware optimization

across GPU, NPU, and edge devices.

Experience in research paper implementation or

open-source contributions.

Preferred Qualifications (Optional):

No

additional preferred qualifications were specified.