Shutterstock: Is Multimodal AI a Data Licensing Problem?

Daniel Mandell is an innovative business leader with over 20 years of experience building and scaling digital media, data and AI-driven businesses.
As SVP of Data Licensing & AI Services at Shutterstock, Dan leads a high-performing team focused on unlocking the value of data through strategic licensing, data services and AI-powered solutions for top AI companies and model builders.
He has a track record of creating and commercialising new ventures, expanding global partnerships and driving transformative growth across media, technology and emerging platforms. He is recognised for combining strategic vision with operational discipline and for building and mentoring high-impact teams.
AI Magazine spoke with Dan about his role the AI space, multimodal data and human creativity in training datasets.
Tell us a bit about you and your role
Iâm the SVP of Data Licensing & AI Services at Shutterstock, where I lead our global efforts across AI, data licensing and AI-driven solutions. My role sits at the intersection of enterprise strategy, partnerships and emerging technology, focused on helping leading AI companies access and operationalise the data they need to build and scale high-performing models.
Iâve spent over 20 years building and growing digital media and technology businesses and whatâs most compelling right now is how rapidly the AI landscape is evolving. The shift isnât just about scale anymore; itâs about precision. Companies need high-quality, rights-cleared and highly specialised datasets that can meaningfully improve model performance and accelerate deployment.
Tell us about Shutterstock's efforts in the AI space
Shutterstock is an end-to-end data partner for AI teams: a single, trusted source where teams can get specialised data quickly and use it confidently, no matter where they are in the model lifecycle.
We offer one of the worldâs largest rights-cleared multimodal datasets, spanning image, video, audio, 3D models, templates, fonts and more. Weâre also always expanding our datasets, which supports ever-evolving model needs. Combined, this enables us to provide fast, centralised access to specialised data required for advanced model development and refinement.
Beyond data, we provide custom data capabilities, including bespoke dataset creation, enrichment and evaluation, to transform raw content into quality training assets. By blending human creative expertise with ML-assisted tools, we can generate structured signals around quality, intent and aesthetics.
Together, this means teams can get exactly what they need, whether off-the-shelf or built to spec, delivered at speed and with the quality and rights confidence required for production AI.
How is the multimodal data bottleneck reshaping AI priorities today?
The biggest constraint for model builders today is immediate access to high-quality, production-ready data across modalities, such as image, video, audio and text.
As models advance, performance gaps are increasingly tied to data: missing alignment between formats, insufficient labelling depth, or lack of real-world context. That shift is pushing the industry to prioritise data quality, specialisation and usability over pure scale.
Itâs also driving demand for high-signal data: data that is intentionally curated, structured and enriched to address specific model needs. Simultaneously, weâre seeing greater demand for human-in-the-loop feedback and evaluation to capture nuance, things like aesthetics, intent and creative quality, that models canât reliably infer on their own. Critically, speed has become a core requirement; teams need the right data quickly and the ability to move from identified need to delivery without friction.
At Shutterstock, this reinforces our focus on delivering specialised, rights-cleared multimodal data and custom data capabilities that help teams fill specific gaps in their training sets and get to deployment faster.
Why is multimodal data harder to source and scale than text?
Multimodal data is inherently more complex. With text, data can often be collected, cleaned and processed in a relatively uniform structure. But multimodal data requires precise alignment across all formats.
Itâs also harder to source at scale because high-quality multimodal content is more expensive to produce and often comes with stricter rights and licensing requirements. Unlike text scraped at scale, multimodal assets require structured creation, deeper metadata and often expert annotation to capture meaning, intent and aesthetic quality.
That's why rights frameworks, labelling depth and accuracy and a trusted sourcing infrastructure aren't optional. They're what makes multimodal data usable for production AI development.
Where does synthetic data fall short for multimodal AI?
Synthetic data doesnât necessarily âfall shortâ as much as it serves a different role. When clearly labelled and responsibly used, it has enormous value for scaling training sets and accelerating development. It can help models learn patterns that would otherwise be expensive, slow, or impractical to capture in the real world.
However, synthetic data will never replace human created data. Multimodal AI depends on authentic human context, nuance and unpredictable interactions across text, image, video and audio that are difficult to fully replicate synthetically. Real-world data provides the grounding models need to understand ambiguity, cultural context and edge cases that emerge naturally in human communication and environments.
What makes multimodal data 'high-quality' beyond sheer volume?
What makes multimodal data âhigh-qualityâ goes far beyond sheer volume. High-quality data is relevant, well-structured and aligned to the real-world use case the model is being trained for. The ability to curate or capture the right data at the right time is what ultimately makes a dataset valuable for AI development.
Quality also depends on strong alignment across modalities. Images, video, audio and text need to connect accurately in time, meaning and intent so models learn how signals relate to one another rather than simply coexist.
Provenance is equally important. Knowing where the data comes from and that itâs responsibly sourced is critical for building reliable, production-grade AI systems.
Finally, fidelity matters. Clear visuals, accurate audio and well-structured metadata all improve signal over noise.
These elements have to work together. Shutterstockâs strength is not just scale, but the ability to source, create, annotate and license multimodal datasets that are relevant, trusted and immediately ready for AI training.
Why is human creative expertise still essential in training datasets?
Multimodal AI doesnât just need to understand what something is, it needs to understand whatâs good. More importantly, it needs to understand why something is good and conversely, why something is bad. That comes down to taste, intent and judgment, which are inherently human qualities. The best creatives act as tastemakers. They know when something resonates, when it feels off, or when itâs technically correct but creatively flat.
Automation can identify errors, inconsistencies, or patterns at scale, but it canât calibrate for taste. It can measure what is common, but not necessarily what is meaningful, distinctive, or emotionally effective. That gap is where human expertise becomes essential. Human creatives bring context, cultural fluency, emotional intelligence and aesthetic judgment into the training process.
Ultimately, that human layer is what helps models move beyond simple pattern recognition and toward a deeper understanding of quality, nuance and impact.


