02 Sep Symposium on Visual Investigations: Best Practice in Collaboration and Community-led Approaches – Using LLMs for Large-scale Video Analysis – Lessons Learnt from VSIN in the Case of the Bangladesh Protest Archive
[Subinoy Mustofi Eron is Design Director at Netra News and a co-founder of Activate Rights, a Bangladesh-based digital rights collective.
Georgia Edwards is Visual Investigations Training & Partnerships Manager at WITNESS.
Dean Issacharoff is the creator of VSIN, an AI powered visual investigations platform.]
Introduction
In July 2024, hundreds of thousands of Bangladeshis took to the streets in an uprising against Sheikh Hasina’s government. In a failed attempt to violently suppress the protests, the government and its supporters killed thousands of people and injured many more. To allow for a public reckoning with the crimes and human rights abuses committed against protesters, WITNESS and Activate Rights created the Bangladesh Protest Archive (BPA). This community-led memorialization initiative collects videos, images, and testimonies from protesters. The archive currently contains more than 10,000 images and videos documenting state violence during the Uprising.
After collecting the archive, activists began to consider how its materials could be activated, through visual investigations that reconstruct key events using open-source information (OSI) verification techniques such as content analysis, AI detection, and geospatial analysis. This presented Activate Rights with a challenge faced by the human rights researchers around the world: making sense of thousands of images and videos of the same events, recorded from different perspectives and shared across social media.
Cataloguing these chaotic media collections is time-consuming and often requires months of manual annotation, as well as repeated exposure to graphic material such as footage of live fire directed at protesters. This was the context in which WITNESS and BPA collaborated with VSIN, a visual investigation platform built to accelerate media analysis.
VSIN helps researchers pinpoint relevant footage and identify patterns involving time, location, and identity. Researchers can build custom analysis workflows based on each file’s content, applying only relevant prompts. The platform also supports automated geolocation, facial and voice matching across videos and images, and multilingual transcription. Once researchers have reviewed the results, they can search, filter, and visualize the media collection to reconstruct events.
Using the VSIN Platform in the Case of “Battleground Uttara”
BPA first used VSIN for “Battleground Uttara,” a forthcoming investigation conducted in collaboration with Netra News, an independent investigative platform in Bangladesh. The investigation reconstructs the violence that took place in Uttara, a district of Dhaka, and examines the tactics used by police and security forces against protesters. VSIN helped researchers find relevant footage faster and curate a focused media collection. The following sections describe the AI and machine learning powered features used during the investigation, how they helped, and their limitations.
The first feature researchers used to process the media collection were deduplication and near-duplicate detection. This is a common step in visual investigations because the same videos are often shared across multiple social media platforms, resulting in different versions entering the archive. Manual deduplication is a tiresome process that requires researchers to review each video and determine which files contain the same or similar footage.
Using open-source machine learning models, VSIN indexes videos according to their audiovisual content. Researchers can then quickly compare indexed media and identify duplicates, including versions that have been clipped or degraded. The results are displayed in a media grid, where VSIN identifies the “original” file based on its duration and quality. Researchers remain in control by reviewing and confirming the model’s results.
Deduplication narrows the media collection, saving time and money during subsequent analysis. Researchers can then build custom workflows using a node-based interface that shows which questions are applied to each file and which models process them. Workflows can branch based on the answers to specific questions. For example, a workflow might ask, “Do uniformed individuals appear in the media?” If the answer is yes, it can ask follow-up questions about their actions and whether they used weapons. Taken together with other analysis, this could support the examination of tactics employed by the state and security forces during the protests.
Descriptions generated by proprietary models such as Gemini are indexed alongside the audiovisual content of each 30-second clip, allowing researchers to search the collection using natural language. For example, searches for “train tracks” or “pink shirt” return timestamps from videos across the archive that most closely match the indexed descriptions and visual content.
Alongside these proprietary models, VSIN uses open-source models for facial and voice matching and object detection. Facial and voice matching can help researchers find recurring appearances across the collection without establishing a person’s identity. Object detection was particularly useful because it allowed researchers to find timestamps where the model detected a weapon, with each detected object marked by a bounding box.
Automated geolocation was another useful feature. Using Gemini’s ability to ground queries in Google’s location services, the model generated likely GPS coordinates and street addresses based on visual clues such as street signs, storefronts, and landmarks. The results were often useful, even when the source material was incomplete or low resolution. In one image from the Barishal Bus Stand, a petrol station called Barishal Auto Service was partially visible. Although the word “Barishal” could not be read in full, the model was still able to identify the location.
Other geolocation results demonstrated how models still lack hyperlocal context and require human verification. For example, the model wrongly claimed that a video was filmed at Sheikh Borhanuddin College because of a misleading sign. The sign actually indicated where the road to the college begins, rather than the location of the college itself. In another example, the model recognized a police box belonging to the Fulbaria Traffic Zone in Chankharpul and incorrectly placed the image in Fulbaria. Chankharpul falls within the Fulbaria Traffic Zone, but the traffic zone and the physical location are not the same. VSIN can generate possible geolocations and surface useful clues, but researchers with knowledge of the local context must verify their reliability.
Keeping Humans in the Loop in Human Rights Investigations
A central question in collaborations like this is how to keep human researchers in control as AI becomes a more powerful tool for media analysis. In VSIN, researchers can review model-generated results and mark them as accurate. This helps maintain the standards required for human rights investigations, particularly when models generate unsupported descriptions.
The quality of model-generated descriptions depends heavily on the context and instructions provided by researchers. For example, a researcher might add the following system prompt to every request: “Act as a human rights researcher documenting the Bangladesh student uprising. Focus on objective descriptions of audiovisual content without assigning judgment.”
A different system prompt could frame the same footage from the perspective of the former government: “Act as an analyst documenting threats to public order during civil unrest. Focus on illegal activity, property damage, and threats faced by security forces.” Given the same files, the model may produce markedly different descriptions. A person described as a “protester” under one framing might be called a “rioter” or “criminal” under another.
To minimize the effect of context on the answer, researchers must know what to ask, focusing on observable details rather than subjective interpretations. A model may identify that someone appears to be handcuffed and arrested, but it cannot reliably determine whether the arrest was unlawful. That conclusion requires legal knowledge, contextual understanding, and human judgment.
This approach was implemented in the BPA investigation. AI was used for the identification of observable details, such as locations, signs, police, weapons, gunshots, smoke, and crowds. These results helped researchers surface relevant footage and flag visual clues, but every finding still required human verification.
This approach is further reflected in the collaborative process used by BPA, WITNESS, and Airwars to develop the BPA Codebook. More information is available in the Opinio Juris article as part of this Symposium, “Data Design and Analysis for Human Rights Audiovisual Archives: Airwars and Bangladesh Protest Archive,” by Clive Vella, Shoeb Abdullah, and Kartika Pratiwi.
Questions, Challenges, and Opportunities
To conclude, we’ve found that the use of AI for large-scale media analysis in human rights investigations can help accelerate the work of researchers, but it also raises important questions about how to preserve human review and agency. This is often described as keeping humans in the loop. Models can help catalogue, search, geolocate, transcribe, and annotate large media collections, but different outputs carry different levels of risk when they are inaccurate. A transcription error may be relatively easy to identify and correct, while an incorrect identification, geolocation, or description of an action could significantly affect the course of an investigation. For “Battleground Uttara,” it was therefore important to establish meaningful points of human review throughout the analysis workflow. Here were some of the questions that emerged related to this:
What Happens when AI-generated Content Becomes a Researcher’s First Impression?
A more subtle question is how generated annotations influence researchers’ first impressions of media. If the first thing a researcher sees when opening a video is a model-generated description, it may shape how they interpret the footage. Even when clearly labelled as generated, the description creates an initial framing, drawing attention toward certain details and away from others. This raises open questions about interface design. Should researchers see the original media before the model’s interpretation? Should generated descriptions remain hidden until requested? How should uncertainty and unverified information be communicated? As discussed in a previous WITNESS report, LLMs can present incorrect information with unwarranted confidence. Asking a model to assess its own certainty does not provide a reliable measure of whether its description is accurate.
How Can AI-surfaced Findings be Documented and Audited Properly for Presentation and Accountability Processes?
This question engages with thinking and organising by Fenix Foundation and Starling Lab, among others, to better understand the implications and necessary practices needed for evidence to be presented in accountability processes that have been analysed or collected using AI. The documentation process and traceability of OSINT practices has been a key strength of the credibility of visual investigations, as well as understanding the provenance and chain of custody of a video for legal proceedings. But due to the opaqueness of certain decisions AI models perform, plus the data that they are trained on – which are used by certain proprietary systems, can there be a documentation process that enables enough explanation to explain the decisions made? Is there a specific model of a hybrid process with the “human in the loop” that can allow for this?
Photo attribution: Pawel Czerwinski on Unsplash

Leave a Reply