Job Description
Your Role:
- Deploy and optimize AI models on edge devices such as mobile, embedded, and IoT platforms.
- Efficiently port and extend open-source inference frameworks like ONNXRuntime, TFLite, etc.
- Build a unified deployment toolchain to enable the principle of "develop once, deploy anywhere".
- Apply lightweight optimization techniques such as pruning and quantization to improve model performance.
- Establish performance evaluation systems to monitor inference speed, resource usage, and key efficiency metrics.
- Write standardized deployment documentation and build reusable component libraries.
✅ Required Qualifications
- Master’s degree or above in Computer Science, Electrical Engineering, Automation, or a related field
- 5+ years of experience in edge computing, embedded systems, or AI deployment
- Strong proficiency in C++17 with solid system-level programming skills
- Familiar with at least two inference engines such as TFLite, MNN, ONNXRuntime, or NCNN
- Hands-on experience with model conversion tools (e.g., tf-lite-converter, ONNX export pipelines)
- Proficient in Linux development environments, including CMake and cross-compilation workflows
- Practical experience deploying models on embedded devices such as Raspberry Pi or NVIDIA Jetson
✨ Preferred Qualifications
- Experience with model quantization (INT8/FP16) and distillation
- Familiarity with edge hardware acceleration (e.g., NPU, GPU)
- Rust development experience
- Understanding of edge-cloud collaborative deployment architectures
- Experience optimizing for ARM architecture, including NEON instruction set
Location:
Silicon Valley (preferred), Suzhou, or Remote
Job Tags