Finding brand marks across advertising video.
Computer vision, language models and speech transcription in one detection pipeline.

The challenge
Brand identification in video is not limited to a static, clearly visible logo. Visual marks, written names and spoken references provide different signals, while manual labelling can be expensive.
The pipeline combines object detection with language-model interpretation and speech transcription. These modalities support logo and brand-name detection in non-English advertisements, with a workflow intended to reduce dependence on extensive manual labels.

Contribution and outputs
Primea developed a multimodal video-analysis pipeline to detect logos and brand names in non-English advertising content.
- Video object-detection pipeline
- Language-model integration
- Speech-transcription integration
- Combined logo and brand-name analysis
Related work
An Arabic interface for specialised investment assistants.
AI Systems
Right-to-left product engineering, persistent conversations and separate knowledge contexts.
Speech interfaces for Pakistani languages.
AI Systems
Recognition, synthesis and voice access to digital information.

