I am an AI engineer and Hong Kong new-media artist — BA in Visual Arts (Studio and Media Arts) and MSc in Data Analytics and AI, both from HKBU's Academy of Visual Arts and Department of Computer Science, respectively, from HKBU, with works shown at the SIGGRAPH Asia 2020 Art Gallery and international exhibitions, and an upcoming part-time lecturer in the Arts and Technology programme at HKBU's Academy of Visual Arts.
My research practice sits exactly where your division does: I build generative AI systems for creative activity and evaluate them with real users. My published and prototyped systems — MetaClues (IEEE MetaCom 2025) and MobiClues (under review, IEEE TALE 2026) — share one design commitment: generative AI as a non-prescriptive co-constructor— a collaborator that scores, suggests and provokes, but never overrides the human's creative intent. My proposed PhD direction carries that commitment from learning environments into creative production tools for visual media.
Creation tools for visual media have become remarkably expressive on the input side and increasingly diverse on the output side. The interpretive channel between them remains thin.
Your ScriptViz [1] lets writers declare which visual attributes are fixed and which are variable. ImaginationVellum [2] goes further, making the entire 2D canvas an active prompt space — spatial arrangement, proximity and freeform strokes map onto prompt tokens and generation parameters, including ControlNet weight and classifier-free guidance.
POET [3] automatically discovers the dimensions along which a text-to-image model is homogeneous, expands them to diversify results, and personalises those expansions from user feedback — helping users reach satisfying results in fewer prompts.
What neither provides is an account, legible to the creator, of how a given result relates to their intent — and to the space they have not explored. ImaginationVellum's provenance graph reports what contributed to an output; POET's preference profile carries feedback from the user to the system. In both cases the system holds rich internal signal — embedding similarities, discovered homogeneity dimensions, provenance chains — and returns almost none of it to the creator in interpretable terms.
Two findings make this consequential.
POET reports that participants' selective attentionshaped what they perceived as diverse — they fixated on terrain, or demographic markers, or lighting, while overlooking other dimensions entirely — and concludes that users' own prompts reinforce normative boundaries independently of what the model can do.
ScriptViz's writers wanted visuals that captured essence and mood rather than exact replicas: an evaluative relationship to the output, not a control problem.
What should a generative creation tool tell a creator about a result, so that the creator revises rather than merely accepts or regenerates — and so that the space they are not exploring becomes visible to them?
I call this the interpretive feedback channel, and I want to design, build and study it. The starting point is my own published mechanism: in MetaClues I designed a multi-faceted relevance score under which high ratings can arise from faithful replication or imaginative divergence, so the metric legitimises multiple creative perspectives instead of enforcing one correct answer. My PhD direction generalises this from a learning environment into creative production tools.
Three research questions, one channel.
What can a system say about a generated result — along which dimensions, in what representation — that a working designer can act on? Dimension-wise interpretation (conceptual, stylistic, compositional) rather than a scalar score or a provenance trace.
How should feedback surface the dimensions a creator has notvaried, without prescribing what they ought to want? This is the direct successor to POET's selective-attention finding: expanding the output space does not guarantee the creator perceives the expansion.
When does interpretive feedback support reflection-in-action, and when does it anchor the creator prematurely? Design fixation is a documented risk of generative creativity tools [4], and a feedback channel is as capable of causing it as of relieving it. This is the evaluation question the direction lives or dies on.
A co-creation tool for graphic or moving-image design — poster, typography, or short-form video ideation — that returns dimension-wise interpretive feedback on each generated result and makes unexplored regions visible. Evaluated with design practitioners and art-and-technology students on process measures (revision behaviour, dimensions explored, fixation) rather than output preference alone. My AVA teaching position gives me direct, ongoing access to exactly this participant population, in Hong Kong, in both Cantonese and English.
BA (Hons) Visual Arts; SIGGRAPH Asia 2020 Art Gallery (two works peer-reviewed); exhibitions in Taipei, Berlin, London, Seoul, Saarbrücken and Hong Kong; residencies at Eaton HK and Chelsea College of Arts (UAL); M+ Museum Hackathon 1st Prize (2020). I do not study creative practice from the outside — I have one.
MetaClues (published, IEEE MetaCom 2025) and MobiClues (under review) are full generative pipelines — image-to-text, reasoning-LLM fusion, text-to-image diffusion, image-to-3D — with user studies (SUS/TAM plus framework-specific instruments, N=19 pilot). I designed both the pedagogical interaction and the evaluation. I was leading the technical team for the system development.
My MSc course (Innovative Lab)'s final project built a modular GUI-based grounding pipeline: a retrained YOLOv8 proposal generator, a fine-tuned CLIP ViT-L/14 retriever and a Qwen2.5-VL-3B LoRA reranker, with out-of-domain LoRA adaptation lifting hit@1 from 14.8% to 60.7% on a production site.
I am transparent that my publication record is not in computer vision — but I have demonstrated that I can fine-tune, adapt and rigorously evaluate vision–language models, which is the technical floor this direction requires. What I bring on top of that floor is rarer: the artist–engineer hybrid.
A. Rao, J.-P. Chou, and M. Agrawala. ScriptViz: A Visualization Tool to Aid Scriptwriting based on a Large Movie Database. UIST '24.
N. Marquardt, A. Roseway, H. Romat, P. Panda, M. Pahud, G. Ramos, S. M. Drucker, A. D. Wilson, K. Hinckley, and N. Riche. ImaginationVellum: Generative-AI Ideation Canvas with Spatial Prompts, Generative Strokes, and Ideation History. UIST '25. doi:10.1145/3746059.3747631.
E. X. Han, A. Q. Zhang, H. Zhu, H. Shen, P. P. Liang, and J. Hsieh. POET: Supporting Prompting Creativity and Personalization with Automated Expansion of Text-to-Image Generation. UIST '25. doi:10.1145/3746059.3747710.
S. Wadinambiarachchi, R. M. Kelly, S. Pareek, Q. Zhou, and E. Velloso. The Effects of Generative AI on Design Fixation and Divergent Thinking. CHI '24.