端到端算法工程师实习生
职位说明
1. Develop quantization, sparsity, pruning, and distillation techniques to enhance the production-level autonomous driving models. 2. Optimize, convert, and deploy autonomous driving models (e.g., ONNX models) that operate efficiently on diverse hardware (GPU, CPU, in-house AI ASIC). 3. Design and implement custom kernels using C++ and CUDA to accelerate model operations and pre/post-processing pipelines. 4. Perform a systematic benchmarking, scaling, and validation of inference performance across various hardware platforms (GPU, CPU, in-house AI ASIC). 5. Collaborate with hardware, compiler, and AI infra engineers to achieve efficient and accurate AI model inference. 1. Familiar with techniques like quantization, sparsity, pruning, and distillation for edge and real-time inference. 2. Strong proficiency in Python and C++, and deep learning frameworks such as PyTorch. 3. Hands-on expertise with CUDA programming, low-level performance profiling, and compiler-level optimization.