Kailai Feng
Education
University College LondonOct. 2026
Incoming Ph.D. Student · Supervisor: Assoc. Prof. Jagmohan Chauhan
Harbin Institute of TechnologyAug. 2023 - Mar. 2026
M.Eng. in Computer Science and Technology · Advisor: Prof. Wangmeng Zuo
Harbin Institute of TechnologyAug. 2019 - Jun. 2023
B.Eng. in Artificial Intelligence
Research Experience
- Developed the first unified on-device model for image generation and editing with a 0.39B backbone; achieved GenEval 0.72 and ImgEdit 4.11, outperforming 12B models such as FLUX, with 3-second generation and editing on an iPhone 17 Pro.
- Unified both tasks through in-context conditioning in latent space; proposed Task-Progressive Joint Pretraining, surpassing separately trained generation- and editing-only models.
- Owned pretraining, editing training, joint training, step distillation, and deployment on iOS and Android devices.
- Developed a training-free artistic typography pipeline to balance visual creativity and glyph legibility; improved both artistry and readability while supporting flexible multi-concept customization.
- Used an LLM and Grounding DINO to split glyphs into deformable Subject and legibility-preserving Surrounding branches; reshaped the Subject with DiffVG and SDS loss, then combined regional interpretation and attention-guided composition for geometry/texture control.
- Owned pipeline design, implementation, experimental validation, ablation studies, and manuscript writing.
- Enabled customized multi-concept generation by composing separately trained heterogeneous single-concept models without joint training, model merging, or extra spatial conditions; outperformed training-based methods in prompt–reference alignment.
- Proposed inference-time Multi-Concept Guidance that adaptively refines visual–text attention and aligns image regions with their concepts to reduce attribute leakage and concept interference.
- Designed and implemented the attention-map module, ran baseline experiments, and contributed to analysis and writing.
- Developed a training-free zero-shot classifier from pretrained text-to-image diffusion models, avoiding task-specific training and image reconstruction; the work was selected as an ICME 2024 Oral.
- Proposed cross-attention semantic and self-attention structural distances under Bayes' rule, then combined them for label–image matching.
- Owned method design, full code implementation, experimental validation, ablation studies, and manuscript writing.
- Developed a zero-shot referring image segmentation framework without paired image–text–mask training data; the generative-only version matched weakly supervised methods, while the full model outperformed competing zero-shot methods.
- Generated fine-grained mask proposals from diffusion attention, then matched them to text with a discriminative model.
- Designed and implemented attention-map extraction, ran baseline experiments, and contributed to analysis and writing.
- Introduced an image-free object detection pipeline requiring neither real training images nor manual annotations; reached about 75% of the same-backbone weakly supervised detector trained on real data.
- Expanded class labels into scene descriptions with a language model, synthesized images, and trained detectors from weak labels.
- Implemented the main training pipeline and baseline experiments, and contributed to experimental analysis and manuscript revision.
Industry Experience
ByteDance, Intelligent Creation LabSep. 2025 - Aug. 2026
Research Intern · Efficient and On-Device Generative AI
- Core DreamLite-v1 researcher: training, paper submission, repo release/maintenance, LoRA/deployment instruction, Diffusers integration.
- Led architecture design, deployment, and comparative experiments across candidate models for DreamLite-v2.
- Led the design, evaluation, and deployment of an on-device diffusion benchmark, including experiments with relevant baselines.
Beijing Century TAL Education Technology Co., Ltd.Apr. 2024 - Sep. 2025
Research Intern · Generative AI and Text-to-SVG Video
- Led the VitaGlyph project from method development and experiments through paper submission.
- Developed a text-to-SVG animation pipeline that uses pretrained diffusion models to guide SVG motion generation.
TikTok, ByteDanceNov. 2022 - Jul. 2023
Machine Learning Engineer Intern · Multi-Task Learning and Video Auto-Moderation
- Owned zero-view pre-recall and feature iteration; proposed multi-path recall, improving overall, violation, and unapproved recall by 3%.
- Designed a multi-scenario recall model with Scenario-Aware Attention and shared/scenario-specific Sub-Experts; within a two-tower architecture, improved violation recall by 2% and unapproved recall by 4%.
- Migrated the production pipeline from TensorFlow to PyTorch with functional and performance parity; six models were launched across multiple regions.
Leadership & Service
- Youth Envoy, ITU Generation Connect2024 - Present
- President, Graduate Student Union, HIT Computing2023 - 2024
- President, HIT Student Union2021 - 2022
Selected Honors
- Samsung Scholarship (0.18%)2024
- Tencent Scholarship (3%)2024
- Special Prize, School Scholarship (20%)2023, 2024
- Outstanding Graduate, HIT (10%)2023
Publication status and author position checked against Google Scholar and linked publisher records on 31 Aug. 2026.