一位曾在英国工作的Google Gemini预训练专家(L7-L8职级)近期加入国内某互联网大厂,标志着本轮争抢Gemini人才的首次成功落地。
According to exclusive information obtained by Leifeng, a Google Gemini pre-training specialist has recently joined a major domestic internet company. The researcher’s rank at Google was likely L7–L8, approaching director level. Having worked long-term in the UK, the hire will primarily focus on pre-training work and will frequently travel between the UK and China for business.
An insider noted that this may mark the first time a domestic giant has “found the right person” in its recent scramble for Gemini talent. Previously, several major companies had recruited individuals from Gemini, but some candidates never touched the core work. Many were either involved only in peripheral pre- or post-training tasks, or held purely managerial roles without oversight of pre-training’s most critical components.
Late last year, domestic foundation model teams significantly shifted their focus toward overseas talent specializing in “pre-training.” An industry insider explained to Leifeng that while Chinese large language models previously prioritized chasing scale—ramping up compute and data—this approach quickly hit the ceiling of scaling laws, making performance gains increasingly difficult. The true differentiators now lie in pre-training strategy: technical route selection, data distribution policies, and the application of first- and second-order optimizers.
Every minor strategic misstep can result in massive waste of time and compute resources. For instance, major tech firms and prominent startups like Kimi and Zhipu have harbored significant internal disagreements over optimizer selection during pre-training, with each company employing vastly different optimization strategies. Such insights on how to improve efficiency and avoid pitfalls are rarely documented in academic papers. While certain optimization methods may appear theoretically sound, determining exactly when and how to adjust them in real-world large model training often requires running complete training cycles. Trade-offs between different training schemes are seldom disclosed in full.
Multiple foundation model experts at major companies told Leifeng that technical route decisions within large model teams are heavily influenced by prior experience. One expert noted that during previous scaling-related projects, they had experimented with various training techniques, but some were ultimately abandoned—not necessarily because they failed technically, but due to assessments of stability, risk, and actual returns. Consequently, overseas laboratories capable of training frontier-level models became the primary targets for recruitment. Gemini emerged as the most typical objective. Some personnel believed Gemini’s robust technical foundation and high talent density, combined with Google’s large-scale organizational structure and frequent adjustments, created ample opportunities for poaching.
At the time, the strategy was straightforward: since domestic pre-training had hit a bottleneck, simply recruit those who had already trained frontier models. However, upon direct engagement, companies discovered that very few individuals actually mastered Gemini’s core pre-training methodologies.
Initially, many recruiters’ first instinct was to target Google’s Silicon Valley headquarters. Given that Google’s core AI operations have long been closely tied to the U.S. tech hub—and that core teams at OpenAI, Anthropic, and Meta are also concentrated in America—the assumption seemed logical. Yet the reality proved different. A Silicon Valley-based AI pre-training researcher told Leifeng that Google deliberately retained its most critical pre-training talent at DeepMind’s UK office, physically and operationally separating them from Silicon Valley’s high-visibility environment to minimize poaching risks.
Meanwhile, a growing number of professionals with DeepMind or Gemini backgrounds have joined Chinese tech giants. From Tian Yonglong’s July onboarding at Tencent, to earlier arrivals at ByteDance such as Wu Yonghui and Qiao Siyuan, and Zhou Hao at Alibaba, the trend is clear. While these resumes carry considerable weight externally, Gemini employees candidly told Leifeng that many of these scientists originated from Silicon Valley offices rather than the UK, meaning they lacked deep expertise in Gemini’s core pre-training methods. As a result, domestic companies remain largely navigating pre-training challenges through trial and error.
This realization has prompted some Chinese firms to recognize that if the goal is to secure DeepMind pre-training experts who hold key methodologies, looking solely at Silicon Valley is insufficient. The remote UK office is where the true talent pool lies. The recent shift in recruitment focus from Silicon Valley to London demonstrates that the targeting strategy is becoming increasingly precise.
(The author continues to track the fierce competition for AI talent among major companies, with a related feature currently in development. Practitioners interested in sharing frontline industry insights are welcome to connect via WeChat: Who123start.)
Original reporting by Leifeng. Unauthorized reproduction is strictly prohibited. See reprint guidelines for details.