Research direction · rev. 2026–07

Generative AI as Co-Constructor: Feedback and Control in Human–AI Creation Tools for Visual Media

PREPARED FOR PROF. ANYI RAODIV. OF ARTS AND MACHINE CREATIVITY, HKUST2027 PhD INTAKE
01

Positioning statement

I am a Hong Kong new-media artist and AI engineer — BA in Visual Arts (Studio and Media Arts) and MSc in Data Analytics and AI, both from HKBU, with works shown at the SIGGRAPH Asia 2020 Art Gallery and international exhibitions, and an incoming part-time lectureship in the Arts and Technology programme at HKBU's Academy of Visual Arts.

My research practice sits exactly where your division does: I build generative AI systems for creative activity and evaluate them with real users. MetaClues (IEEE MetaCom 2025) and MobiClues (under review, IEEE TALE 2026) share one design commitment: generative AI as a non-prescriptive co-constructor— a collaborator that scores, suggests and provokes, but never overrides the human's creative intent. My proposed PhD direction carries that commitment from learning environments into creative production tools for visual media.

02

Proposed direction

Core question

When a generative system participates in a creative workflow, how should its feedback and control mechanisms be designed so that it expandsthe creator's space of intent rather than substituting for it — and how do we evaluate that on the creative process, not just the output?

Two anchors, both grounded in your work.

A

Interpretive feedback as a creative signal

In MetaClues/MobiClues I designed a multi-faceted relevance score in which high ratings can come from faithful replication or imaginative divergence — the metric deliberately legitimizes multiple creative perspectives instead of enforcing a single correct answer. I want to develop this into a general mechanism for creation tools: LLM/VLM-based interpretive feedback that tells a creator how their draft relates to a reference, a brief, or a style — conceptually, culturally, visually — as a provocation for iteration.

This extends the evaluative dimension your ScriptViz user study surfaced — writers wanting visuals that capture essence and mood rather than exact replicas — into an explicit, designable feedback channel.

B

Structured control between consistency and exploration

ScriptViz's fixed/variable attribute control demonstrates that creators benefit from declaring what is settled and what is open. I see a direct line from this to controllable generation — AnimateDiff's motion priors, VideoRepainter's keyframe-then-propagate decomposition: the HCI question is how to surface model-level control (conditioning, keyframes, LoRA-adapted styles) as creator-level commitments about what stays fixed and what may vary. My interest is in building and studying tools that make this negotiation legible for working designers and media artists.

Candidate first project — illustrative, not fixed

A co-creation tool for graphic and moving-image design — poster, typography, or short-form video ideation — combining Anchor A's interpretive scoring with Anchor B's fixed/variable control, evaluated in a user study with design practitioners and art-and-technology students. My AVA teaching position gives me direct, ongoing access to exactly this participant population, in Hong Kong, in both Cantonese and English.

03

Evidence of capability

As an artist

BA (Hons) Visual Arts; SIGGRAPH Asia 2020 Art Gallery (two works peer-reviewed); exhibitions in Taipei, Berlin, London, Seoul, Saarbrücken and Hong Kong; residencies at Eaton HK and Chelsea College of Arts (UAL); M+ Museum Hackathon 1st Prize (2020). I do not study creative practice from the outside — I have one.

As a builder of human–AI creative systems

MetaClues (published, IEEE MetaCom 2025) and MobiClues (under review) are full generative pipelines — image-to-text, reasoning-LLM fusion, text-to-image diffusion, image-to-3D — with user studies (SUS/TAM plus framework-specific instruments, N=19 pilot). I designed both the interaction and the evaluation.

As a vision–language practitioner

My MSc final project built a modular GUI-grounding pipeline: a retrained YOLOv8 proposal generator, a fine-tuned CLIP ViT-L/14 retriever, and a Qwen2.5-VL-3B LoRA reranker, with out-of-domain LoRA adaptation lifting hit@1 substantially on a production site.

CLIP retriever · hit@20
0.214 → 0.554
LoRA adaptation · hit@1
14.8% → 60.7%

I am transparent that my publication record is not in computer vision — but I have demonstrated that I can fine-tune, adapt and rigorously evaluate vision-language models, which is the technical floor this direction requires. What I bring on top of that floor is rarer: the artist–engineer hybrid.

04

Why your group

Your division exists to merge art and technology; your lab recruits artists alongside researchers; your own trajectory runs from model contributions (AnimateDiff, IC-Light, MovieNet) to human-centered tools with serious user evaluation (ScriptViz, UIST '24; CineVision, UIST '25).

My profile mirrors the tool-and-study end of that spectrum with enough model-level competence to prototype independently — and adds a practicing artist's perspective, Hong Kong arts-community ties through HKBU AVA and the local exhibition network, and teaching experience in exactly the creative-technology pedagogy your division delivers.