Our Work / Application development

Finding brand marks across advertising video.

Computer vision, language models and speech transcription in one detection pipeline.

Multimodal logo detection — conceptual artwork in graphite, silver and burgundy.

The challenge

Brand identification in video is not limited to a static, clearly visible logo. Visual marks, written names and spoken references provide different signals, while manual labelling can be expensive.

The pipeline combines object detection with language-model interpretation and speech transcription. These modalities support logo and brand-name detection in non-English advertisements, with a workflow intended to reduce dependence on extensive manual labels.

Explanatory advertising-video analysis diagram: object detection produces visual candidates, speech transcription provides spoken references, and written-name evidence joins them for combined review with language-model interpretation.

Contribution and outputs

Primea developed a multimodal video-analysis pipeline to detect logos and brand names in non-English advertising content.

  • Video object-detection pipeline
  • Language-model integration
  • Speech-transcription integration
  • Combined logo and brand-name analysis

Related work

Bring us the hard part.

Discuss your challenge