blackforestlabs

Member of Technical Staff - VLM

Full-timeNot specifiedRemoteRemoters,Apply by 12 Oct 2026
Apply now
WA

Overview

About Black Forest Labs We're the team behind Latent Diffusion, Stable Diffusion, and FLUX, foundational technologies that changed how the world creates images and video. Our models power the tools used by millions of creators, developers, and businesses worldwide, and FLUX is among the most advanced generative systems in the world. Headquartered in Freiburg, Germany with a growing presence in San

What you'll do

  • About Black Forest Labs We're the team behind Latent Diffusion, Stable Diffusion, and FLUX, foundational technologies that changed how the world creates images and video.
  • Our models power the tools used by millions of creators, developers, and businesses worldwide, and FLUX is among the most advanced generative systems in the world.
  • Headquartered in Freiburg, Germany with a growing presence in San Francisco, we're scaling fast while staying true to what makes us different: research excellence, open science, and building technology that expands human creativity.
  • Why This Role Vision-language models are becoming foundational to how people interact with generative AI, but most VLM research happens in isolation from the generation stack.
  • At Black Forest Labs, we're integrating VLMs directly into FLUX in ways that make our models more powerful, more controllable, and more aligned with what creators actually want.
  • This role is about pioneering that integration.
  • You won't be applying off-the-shelf VLMs, you'll develop novel approaches, innovate on architectures, and answer questions that haven't been solved yet: how vision and language representations inform each other, how multimodal understanding improves generation quality, and how to make these capabilities deployable at scale without compromising what makes FLUX exceptional.
  • This is a Staff / Senior IC role.
  • We're looking for someone who has pretrained or significantly advanced a VLM, not just fine-tuned one.
  • What You'll Work On Lead development and training of state-of-the-art multimodal vision-language models within the FLUX stack, innovating on architectures, not just applying existing ones Design fine-tuning strategies that adapt VLMs to specialized creative use cases (captioning, editing instructions, prompt enhancement) that general-purpose models can't handle Research integrations between VLM/LLM capabilities and our diffusion and flow pipelines, finding creative ways to improve generation quality and controllability without computational bottlenecks Evaluate emerging multimodal architectures, translating the best of recent research into practical improvements What We're Looking For You've pretrained or significantly advanced a VLM (not just SFT'd or LoRA'd one) that was deployed in a production system or released publicly Strong publication record or unambiguous production track record showing you push.

Requirements

  • Practical experience with Research.
  • Relevant academic or project background for a Member of Technical Staff - VLM role.
  • Strong written and verbal communication in English.
  • Comfortable working remotely with distributed teams.

Skills

Research

Related roles