招聘信息

Forward Deployed Engineer (Inference & Post-Training) - Mandarin Speaking

Key takeaway: 本文是 Together AI 通过 Zero G Talent 发布的一份面向新加坡的招聘启事,岗位为会说普通话的前线部署工程师(Forward Deployed Engineer,FDE),专门负责推理(Inference)与后训练(Post-Training)方向。该职位不是简单的解决方案架构师替代品,而是定位于深度领域专家,需要与解决方案架构师协同工作。核心职责围绕推理引擎优化展开,涉及根据硬件配置、模型架构和负载特征选型、配置并优化主流推理引擎(如 vLLM、TensorRT-LLM、SGLang)。具体技术工作包括:调整 KV Cache、应用推测解码、确定最佳张量并行度与量化策略,以达成严苛的吞吐量和延迟指标。后训练方面,候选人需亲自执行 RL 训练任务并优化系统架构,指导客户完成 LoRA、SFT、DPO、RLHF 与 GRPO 等全流程管线搭建,助力客户从实验阶段跨入生产环境。此外,FDE 还需负责战略客户的长期技术健康度,确立客户入驻平台的基线配置以缩短价值实现时间,并代表一线经验反向驱动产品路线图的演进。申请者须具备 5 年以上技术经验,对开源大模型生态有广泛认知,具备专家级推理引擎实操和诊断能力,并拥有扎实的 Python 编码功底。文章强调 Together AI 是一家研究驱动型企业,致力于通过软硬件协同设计降低 AI 成本,其团队贡献了 FlashAttention 等知名技术。岗位要求为新加坡永久居民或公民,提供初创股权及远程灵活办公选项,发布时间为 2026 年 8 月 4 日。

Key Takeaways
  1. Together AI 正在新加坡招聘专注于推理与后训练的前线部署工程师(FDE),要求具备流利普通话能力及新加坡永居/公民身份。
  2. 该职位区别于传统的解决方案架构师,要求深入参与客户 POC 验证、推理引擎选型、性能调优及后训练流程指导,直接对客户成功和平台迭代负责。
  3. 工作要求具备 vLLM、TensorRT-LLM、SGLang 等推理引擎的专家级实操经验,并能针对性进行 KV Cache 调优、推测解码、张量并行及量化策略调整。
  4. 需熟练掌握 LoRA、SFT、DPO、RLHF 及 GRPO 等后训练与微调流程,辅助客户完成从实验到生产环境的全链路部署。
  5. Together AI 强调以透明开放的方式降低 AI 系统成本,其团队贡献了 FlashAttention、Hyena、FlexGen 和 RedPajama 等知名技术成果。
这是一则来自 Together AI 新加坡团队的重量级招聘信息,精准定位了当下大模型落地最稀缺的实战型角色——前线部署工程师(FDE)。与传统的解决方案架构师不同,该岗位要求候选人同时具备推理引擎源码级的调优能力和后训练全流程的实战经验。透过这份 JD,从业者可以清晰看到头部 AI 公司对 vLLM、SGLang 等推理框架以及 LoRA、DPO、RLHF 等微调技术的硬性要求。对于希望向大模型工程化方向转型的技术人员而言,这份招聘信息本身就是一份极具价值的“技能图谱”。

Together AI 正在新加坡招聘专注于推理与后训练的前线部署工程师(FDE),要求具备流利普通话能力及新加坡永居/公民身份。

—— 络石智能编辑部 · Editor's Pick

Forward Deployed Engineer (Inference & Post-Training) - Mandarin Speaking

Together AI•

Singapore, Singapore

Job Description

About the role

As a Forward Deployed Engineer (FDE) focused on Inference & Post-Training, you will be a hands-on technical partner to our most strategic customers — production AI teams looking to leverage high quality models and do inference at scale. For us, FDE is not a replacement for a Solutions Architect; you will partner with our SAs as a deep-domain specialist in inference optimization, fine-tuning pipelines, and production deployment. As key contributors to both the CX, Engineering, and Sales organizations, FDEs add tremendous value by ensuring we can meet the requirements of our most complex POCs, facilitate successful platform adoption, and guide tailored optimization efforts — directly impacting customer success, company growth, and the hardening of our core platform.

Must be a permanent resident or citizen of Singapore.

Responsibilities

  • Inference Engine Optimization: Select, configure, and optimize inference engine based on hardware, model architecture, and workload profile
  • Configuration & Performance Tuning: Develop configuration updates to win critical POCs, benchmarks, and optimize customer deployments; tune KV cache, apply speculative decoding, determine optimal tensor parallelism, and determine quantization strategy to hit throughput and latency targets.
  • Post-Training & Fine-Tuning: Drive hands-on RL training runs and optimize system design; guide customers through LoRA, SFT, DPO, RLHF, and GRPO pipelines from experimentation through production.
  • Strategic Customer Alignment: Act as the primary technical point of contact for aligned strategic accounts — monitoring and optimizing endpoint configurations, helping customers get the most out of the platform, and collaborating to ensure we hit critical milestones.
  • Opinionated Onboarding: Establish direct alignment with strategic customers at onboarding; ensure the right inference and post-training configurations are in place from day one to improve time-to-value.
  • Product Feedback Loop: Directly influence our software and model roadmap by surfacing insights from the field. Contribute back to the product where needed to support customer requirements or drive a better experience. Drive early feature and research adoption with strategic logos.

Qualifications

  • Experience: 5+ years in a technical role, with a strong focus on inference systems, open-source LLM deployment, or post-training workflows.
  • Inference Engine Depth: Expert-level, hands-on experience with inference engines (e.g., vLLM, TensorRT-LLM, SGLang); ability to diagnose and resolve performance issues at the engine level.
  • Inference Optimization: Deep knowledge of KV cache tuning, speculative decoding, tensor parallelism, pipeline parallelism, and quantization techniques
  • Post-Training Knowledge: Hands-on experience with fine-tuning and post-training pipelines, including LoRA, SFT, DPO, RLHF, and GRPO; ability to advise on system design
  • Model Landscape Awareness: Broad knowledge of state-of-the-art open-source models and strong judgment on model selection for specific customer use cases, hardware profiles, and performance targets.
  • Coding Proficiency: Strong Python skills; comfortable working in production environments

About Together AI

Together AI is a research-driven artificial intelligence company. We believe open and transparent AI systems will drive innovation and create the best outcomes for society, and together we are on a mission to significantly lower the cost of modern AI systems by co-designing software, hardware, algorithms, and models. We have contributed to leading open-source research, models, and datasets to advance the frontier of AI, and our team has been behind technological advancements such as FlashAttention, Hyena, FlexGen, and RedPajama. We invite you to join a passionate group of researchers on our journey in building the next generation of AI infrastructure.

Compensation

We offer competitive compensation, startup equity, health insurance, and other benefits, as well as flexibility in terms of remote work. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge.

Equal Opportunity

Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.

Tags

Related Topics

Expert Comment

This article is curated by the editorial team from public sources for reference only.

FAQ

Together AI 新加坡正在招聘的前线部署工程师(FDE)核心职责是什么?其与解决方案架构师有何不同?
Together AI 的前线部署工程师(FDE)在推理和后训练方面扮演着深度技术合作伙伴的角色,这不同于传统的解决方案架构师。FDE 主要负责推理引擎(如 vLLM、TensorRT-LLM、SGLang)的配置与极致性能调优(包括 KV Cache 调优、推测解码、张量并行和量化策略)。此外,他们需要指导客户实施 LoRA、SFT、DPO、RLHF 和 GRPO 等微调及后训练流程,并确保客户在平台入驻首日就获得正确的技术配置以缩短价值实现时间。该岗位还承担着将一线需求反馈至产品路线图的责任。
申请 Together AI 的 FDE 职位需要具备哪些具体的技术栈与经验门槛?
基础资格包括 5 年以上技术经验,且需深耕推理系统或大模型部署。硬性技能树涵盖了专家级的推理引擎操作能力(vLLM/SGLang/TensorRT-LLM),深谙各种推理优化手段(如 KV Cache、量化);同时必须具备 LoRA、SFT、DPO、RLHF、GRPO 的实战后训练经验。软性要求则包括对主流开源模型生态的广度认知,出色的 Python 编程功底,以及直接对齐战略客户需求的能力。申请人必须是新加坡永久居民或公民。

Related Articles

Handshake 高级 AI 前部署工程师

本文是Handshake公司发布的“高级AI前线部署工程师”(Senior AI Forward Deployed Engineer)招聘启事。Handshake从大学生求职平台转型为AI数据服务巨头,在2025年启动Handshake AI业务,号称已成为史上增长最快的AI数据企业,年化营收约10亿美元,每月向超过3万名知识工作者支付约6,000万美元。该职位位于旧金山,采用混合办公模式,核心使命是嵌入前沿AI实验室等战略合作伙伴中,扮演应用AI研究与客户交付的桥梁角色。工程师需要将实验室模糊的后训练需求转化为具体的评估框架、标注管道和基准基础设施,并主导原型设计、实验迭代和团队指导。技术要求方面,候选人需有6年以上应用机器学习或AI研究工程经验,深度掌握强化学习后训练技术(RLHF、DPO、PPO),具备实际模型微调操作经验(LoRA、PEFT),并有ML数据管道和评估工具的使用背景。加分项包括LLM评估设计经验、已发表的研究成果、开源贡献或在高增长AI公司担任过前线部署类角色。该岗位旨在影响前沿模型的训练方式,反映了AI数据服务行业对兼具研究深度与工程交付能力的复合型人才的迫切需求。

Read More

AI现场工程师 - GenAI基础设施

2026年7月8日,一家匿名的 GenAI 基础设施公司(Stealth)发布招聘信息,寻找 AI 前线部署工程师(AI Field Engineer)。该职位旨在帮助企业将开源大模型从沙盒探索推向生产环境,职责包括确定概念验证范围、运行负载测试、执行 SFT/DPO/RFT 等模型微调,并直接在客户基础设施中交付生产集成,而非停留在咨询层面。公司推理平台已为 Uber、DoorDash、Notion 和 Cursor 等知名企业提供生产服务,推理速度达到闭源模型的 15 倍,并发量提升 4 倍。职位要求候选人具备 3 年以上客户现场 AI/ML 工程经验,深度掌握 LLM 推理与训练,熟练使用 vLLM、SGLang 或 TensorRT-LLM 等推理引擎,精通 Python 编程及 AWS、Azure、GCP 等 GPU 云基础设施,并具备 Kubernetes 实战能力。理想候选人拥有 Palantir FDE、BCG X 或 McKinsey QuantumBlack 等前线部署或专业服务背景。薪资范围为底薪 176,000–224,000 美元,总包 220,000–280,000 美元(80/20 拆分),10 年以上经验可上浮,并提供股权激励。工作方式为美国境内远程办公,需定期出差至客户现场,并支持 H-1B 转移和 TN 签证担保,O-1 签证视情况而定。该招聘突出强调了实战交付能力,明确拒绝仅有闭源 API 经验的候选人,反映出开源模型生产部署对 FDE 角色的硬核要求。

Read More

腾讯混元AI Infra如何优化Hy3 Preview:一次大模型推理性能提升的技术拆解

本文详细介绍了腾讯混元AI Infra推理团队针对旗舰大模型Hy3 Preview(采用GQA+MoE混合架构,原生支持256K超长上下文)在NVIDIA Hopper卡上的推理性能优化实践。面对Hopper卡算力较低、显存紧凑等限制,团队从算子优化与融合、并行策略、多级缓存、MTP异步调度、量化与稀疏五大维度进行全栈优化。在算子优化上,提出动态调度负载均衡的Attention算子(混合长度batch加速1.59x-1.76x),双BF16重构FP32 Router GEMM(加速2.86x-3.22x),FusedMoE流水线重构(相比vLLM等加速1.2x-1.6x)。算子融合方面,实现Fused Rope+Norm+Quant+Store KV(加速约5x),Fused AllReduce+Norm+Add(加速1.68x),采样融合算子(加速2.5x-5.5x),以及Gemm+Comm通算融合(加速1.68x-1.81x)。并行策略上,采用TPSP Prefill优化(TTFT降低24.5%-29.9%)和DP+EP Decode架构(吞吐提升15.7%-44.7%)。多级缓存构建GPU-CPU-KVStore三级体系以降低重复Prefill。MTP异步调度优化消除CPU气泡,端到端提升10%-20%。量化方面,在AngelSlim框架中通过GPTQ权重重建、激活平滑、Hadamard旋转和QAT微调实现W8A8C8无损量化,吞吐提升28%+;并应用Stem稀疏注意力算法及HPC-BSA算子,在128K上下文下Prefill延迟降低3.6倍,精度持平。文章为Hopper架构下大模型推理部署提供了系统级优化范本。

Read More

Qwen3-32B本地部署实战:LoRA微调+vLLM推理+Docker交付

本文是一份针对阿里巴巴通义千问系列开源大模型 Qwen3-32B 的本地化工程部署与微调实战指南。文章围绕“LoRA 微调 + vLLM 推理 + Docker 容器化交付”这一技术路径,详细记录了从环境准备、模型下载与格式转换、Docker 镜像构建、vLLM 高性能推理服务上线、到基于 Swift 框架的 LoRA 领域微调及效果验证的完整流程。核心观点强调 Qwen3-32B 作为“甜点规模”模型,在 A100-80G 单卡上即可完成推理,双卡可实现高吞吐服务,且 MMLU、CMMLU 等评测稳居开源第一梯队。技术选型深度对比了 vLLM 与 TGI 的架构差异,指出 PagedAttention 机制使 32K 长文本推理显存占用比 TGI 低 58%,吞吐量高出 2.8 倍。在 LoRA 配置上,通过梯度分析锁定 q_proj 和 v_proj 为核心目标模块,并给出 balanced 方案(q_proj+v_proj+gate_proj+down_proj,rank=8)在显存占用 31.5GB 下实现 ROUGE-L 提升 4.1 的实测数据。文章还重点描述了 Docker 从零构建的可复现策略、应对冷启动的预热技巧、QWEN 官方 AWQ 量化加速方案,以及显存溢出、模型加载失败等血泪教训。本文面向面临模型本地化、业务定制与生产环境落地的工程师,提供了可复现、可迭代、可交付的完整工作流,具有极强的工程参考价值。

Read More

FDE前沿部署工程师实战指南:从模型部署到AI Agent系统构建

本文是一份面向前沿部署工程师(FDE)的完整实战指南,系统拆解了 FDE 的核心内涵、技能树、项目实战和学习路径。文章指出 FDE 不是简单的“模型部署”,而是确保复杂 AI 能力(大语言模型、多模态模型、AI Agent、RAG)能够在真实生产系统中稳定、高效、可扩展且低成本地集成的全栈角色,其 40k 以上月薪是对“AI 应用最后一公里”复杂性的定价。核心技能体系覆盖 AI 基础(Transformer、Prompt 工程、RAG、Agent)、工程开发(Python、FastAPI 异步编程)、云原生(Docker、Kubernetes、Istio、Prometheus/Grafana)和系统设计。实战部分通过使用 vLLM 部署 Qwen2.5-7B-Instruct-AWQ 量化模型、构建 FastAPI 代理网关实现路由与监控、编写 Dockerfile 和 Kubernetes Deployment 文件完成容器化与编排,以及部署多智能体项目 My AI Town 来串联所有技能。文章还提供了常见故障排查清单(超时、OOM、推理慢、Agent 循环、Pod 重启)、最佳实践(IaC、配置分离、可观测性、成本优化、降级策略)及三阶段学习路径(基础巩固、核心技能、项目实战),并解析了面试考察的系统设计、故障排查和工程协作等方向,为求职者提供了从零到就业的路线图。

Read More