Skip to main content Skip to side menu Skip to main menu

Computer Science

Events

CS Seminar by Alex Bogdan from Avenga & Google X Quest (Friday, October 2, 2 PM, @B206)

Writer Computer Science Date Created 2026.09.21 Hits40

Alex Bogdan from Avenga and Google X Quest will be giving a talk on "Pretrained Vision Transformers Generalize to Non-RGB Arbitrary-Channel Imagery Classification".

Please find the seminar details below:
Title: Pretrained Vision Transformers Generalize to Non-RGB Arbitrary-Channel Imagery Classification
Date & Time: Friday, October 2, 2 PM
Venue: B206

Abstract
Open-weights Vision Transformers (ViTs), pre-trained on massive image datasets, have proven highly effective for transfer learning, especially in the few-shot regime. Despite their success, their application is largely confined to 3-channel RGB inputs.
As a result, fields like medical imaging, remote sensing, and geoscience, which rely on multispectral, hyperspectral, or 3D volumetric data, do not benefit from this common resource because they exceed the standard 3-channel constraint.
To address this shortcoming, we present a simple and compute-efficient recipe for adapting any pre-trained ViT to accept inputs with an arbitrary number of channels, thereby enabling transfer learning. Our early-fusion approach introduces two key concepts: 1) the use of attention pooling across channels after the first convolutional layer, and 2) weight replication of the first convolutional layer, with or without sharing weights. Applied to classification tasks, this approach allows the aforementioned specialized domains to leverage powerful foundational models (e.g. DINO, SigLIP or CLIP variants). Furthermore, we show that this approach is equally effective for fine-tuning on 3D volumetric data. Our approach produces strong baselines which we demonstrate using 4 datasets from diverse domains.

Speaker Bio
Alex Bogdan’s work focuses on computer vision, deep learning, and self-supervised learning, with an emphasis on foundation model adaptation, multimodal sensor pipelines, and edge-deployable detection systems.
He earned his Bachelor of Science in Mathematical Sciences from the Korea Advanced Institute of Science and Technology (KAIST). Over more than seven years in machine learning engineering, he has translated applied research into scalable production systems across industries ranging from robotics to geophysical sensing.
Currently, he works as a Machine Learning Engineer with Avenga, collaborating as an extended workforce contributor with the Google X Quest team on state-of-the-art Vision Transformer (ViT) and DINO architectures for seismic signal classification, volumetric data modeling, and fine-grained segmentation. His industry experience also includes developing LiDAR-based foreign object debris detection for airport runways with ADB SAFEGATE, advanced driver assistance systems (ADAS) at Pittasoft, and tracking frameworks for biomedical robotics. He is also the co-author of multiple peer-reviewed publications in computer vision.