At some point in nearly every enterprise sales conversation we have, someone asks the question: “Couldn’t we just build this ourselves?”
It’s a fair question. The raw ingredients seem accessible: social listening APIs, review scrapers and services, a frontier LLM, and an eager data science team. A working demo can materialise in weeks. And that demo will look great in front of your leadership team.
But a demo is not a product, and a product is not a durable insights capability. The research on internal AI builds is sobering: RAND Corporation found that more than 80% of AI projects fail to deliver their intended business value – roughly twice the failure rate of IT projects that don’t involve AI. MIT’s Project NANDA study of enterprise generative AI found that 95% of pilots deliver no measurable P&L return, and, most relevant to this decision, that externally sourced AI solutions reach successful deployment roughly twice as often as internal builds (~67% vs. ~33%). Gartner reports that at least half of GenAI projects are abandoned after the proof of concept, and S&P Global found that 42% of companies abandoned most of their AI initiatives in 2025, up from 17% the year before.
We explored this exact tension in our recent webinar, Buy vs. Build: Strategic Decision-Making for Future-Ready Organisations, featuring James Cummings, Vijay Raj, and our CEO, Trevor Sumner. If you’re weighing this decision, watch the full discussion – and work through the 20 questions below with your team first. To paraphrase Vijay Raj: you may think you understand the purchase price, but you are signing up for a mortgage of maintenance that can bankrupt you.
The data problem and the feeds that break at 2 a.m. (Questions 1-5)

1. Who owns the data pipeline when it breaks at 2 a.m.?
Consumer signal data is not a static asset – it’s a living, hostile environment. Review platforms change their page structures. Social networks throttle, deprecate, or reprice their APIs (sometimes overnight, as X and Reddit both demonstrated this year). E-commerce sites deploy anti-bot measures or severely limit reviews without a login, as Amazon did this year as well. Every one of these events breaks pipelines, and every break creates silent data gaps that corrupt your trendlines.
At i-Genie.ai, we maintain data partnerships and proprietary collection infrastructure across dozens of platforms as a full-time, always-on function. When you build, that pager belongs to you.
2. What happens to your historical benchmarks when a data source changes?
Brand equity tracking and product benchmarking depend on longitudinal consistency. When a source changes its sampling, format, or availability, someone has to normalise the old data against the new – or your year-over-year comparisons become fiction. This is unglamorous, ongoing statistical work that internal teams consistently underestimate. Who makes that call, who do they need to consult, and how is the work resourced?
3. How will you evaluate and onboard new data sources as consumer behaviour shifts?
Three years ago, nobody was asking about visibility in LLM-powered search. Today, how your brand appears in ChatGPT, Gemini, and Perplexity answers is a board-level question. TikTok Shop reviews barely existed as a signal source; now they’re essential for beauty and CPG categories. Tomorrow it will be something else. i-Genie.ai is already supporting these “tomorrows” while an internal team is still drafting its resource request.
A vendor evaluates, tests, and integrates new sources across its entire client base, optimising for metadata completeness and quality and system integrity. An internal team has to spot the shift, make the business case, fund the integration, and validate the data – every single time – while doing everything else on this list.
4. Do you have the rights, compliance posture and audits to collect this data?
Web data collection sits in an evolving legal and regulatory landscape: platform terms of service, GDPR and its global cousins, PII obfuscation, copyright questions around scraped content, and emerging AI-specific regulation. Vendors amortise legal review, security certifications, specialised insurance and privacy audits across their entire client base. When you build, you carry 100% of those fixed costs – and 100% of the exposure.
5. What does your data actually cost at full scale
Enterprise-grade social listening feeds, review data, search data, and e-commerce data carry meaningful licensing costs – often six to seven figures annually before you’ve written a line of code. Internal builds routinely scope the engineering and forget that the raw data itself is one of the largest line items. We know because we manage these vendor relationships at scale, and we negotiate volume pricing an individual brand can’t access.
The model probem. Keeping pace with a landscape that moves in weeks (Questions 6-11)

6. Who evaluates every new model release and against what benchmark?
The model landscape now moves in weeks, not years. Each major release forces a question: is it better for your specific tasks – sentiment classification, theme extraction, trend detection, multilingual review analysis? Answering that requires maintained evaluation sets with ground-truth labels across every function and language you operate in. Building and maintaining those eval suites is a discipline unto itself. Skipping it means flying blind.
7. Are you continuously optimising the cost/power tradeoff per task?
Not every job needs a frontier model. Tagging a million reviews for basic sentiment is a job for a small, cheap model; nuanced cross-market trend synthesis may justify the most powerful one available. Getting this right requires continuous cost-optimisation analysis: routing each task to the cheapest model that meets the accuracy bar, and re-running that analysis every time pricing or capability changes.
Getting it wrong is expensive in both directions. Route everything to a frontier model and your inference bill explodes – unexpected AI cost overruns are one of the reasons Gartner cited for GenAI project abandonment. In 2026, Uber burned through its entire annual AI coding-tools budget in just four months as usage surged. Who’s accountable for the cost overrun or the system shutdowns? On the other hand, route everything to a cheap model and your insights quietly degrade.
8. What’s your plan when the model you depend on gets deprecated?
Providers regularly sunset older, cheaper models. When that happens, you don’t get a vote – you’re forced to migrate, re-run your evaluations, re-tune your prompts and pipelines, and re-validate outputs against historical baselines, all on someone else’s timeline. Vendors absorb this churn as routine operations. For an internal team, every deprecation is an unplanned project – and an unplanned cost (development, testing, benchmarking, etc.).
9. Have you validated model performance in every language you sell in?
A model that performs beautifully on English-language Amazon reviews may stumble on Bahasa Indonesia social posts, Arabic dialect variation, or Japanese review conventions. Multilingual accuracy isn’t a checkbox; it’s per-language evaluation, tuning, and often per-language model selection. If you operate in 10+ global markets, multiply your entire evaluation burden accordingly. We operate in over 30 countries and 20+ languages. Who do you have on staff that can optimise Tagalog, Thai, Arabic or the numerous regional Indian languages you need for representation?
10. If you fine-tune custom models, where does your training data come from?
Custom models require labelled training data, which means data tagging operations, annotation guidelines, quality control, and crucially, outcome data linking signals to real business results. Assembling this is expensive and slow, and it’s exactly the kind of asset a specialised vendor has been compounding across many brands and categories for years. Worse still: will your company even permit training on first-party data without strict AI governance in place?
11. And what happens when your training data goes stale?
Consumer language evolves. Platforms rise and fall. Data sources you trained on become unavailable. Custom models need retraining on refreshed data – a recurring cost, not a one-time investment. RAND’s research identified inadequate data and infrastructure as core root causes of AI project failure; stale training data is how that failure arrives slowly, then suddenly.
The software and cost problem. What the demo never shows you (Questions 12-16)

12. Have you budgeted for the whole team – not just the ML engineers?
A real insights platform needs data engineers, ML engineers, backend and frontend developers, designers, product managers, QA, DevOps, security, and ongoing user support. How many humans in the loop are you accounting for? With median ML/AI compensation now around $244,500 and AI skills commanding a 56% wage premium (PwC), a credible internal team is a multi-million-dollar annual commitment before infrastructure, data, and inference costs. Does your company pay above market or offer generous equity to attract top talent?
13. What does the history of internal IT projects tell you about your timeline?
The McKinsey-Oxford study of 5,400 large IT projects found they run 45% over budget and 7% over time on average, while delivering 56% less value than predicted – and software projects carry the highest overrun risk. Most alarming: 17% of large IT projects become “black swans” with overruns severe enough to threaten the company itself. Every additional year of project duration adds roughly 15% to cost overruns. If your build plan says 12 months, plan for the plan to be wrong. What’s the cost of time when your markets move faster than ever?
“The real tension isn’t just build vs buy, it’s organisational speed versus market speed.”
Trevor Sumner, CEO, i-Genie.ai.
Buy vs. Build: Strategic Decision-Making for Future-Ready Organisations webinar
14. By the time you’ve built to where vendors are today, where will vendors be?
This is the treadmill problem, and it’s the one internal builds almost never model honestly. It might take 18-24 months to replicate a vendor’s current capability. During those 18-24 months, the vendor with a dedicated team, compounding data assets, and roadmap input from dozens of enterprise clients keeps moving. You’re not building toward the finish line; you’re building toward where the finish line used to be.
“An 8-month internal build project can be completely outpaced by a market-ready product that launches in the same window.”
Vijay Raj, Ex-EVP of CMI, Unilever.
Buy vs. Build: Strategic Decision-Making for Future-Ready Organisations webinar
As we put it in our webinar: this isn’t really a buy-versus-build decision anymore – it’s about how organisations combine speed, control, and intelligence across a much more complex AI ecosystem.
15. Who handles the feature request backlog after launch?
The day your internal platform launches, it becomes an internal product with internal customers and a backlog of years of work. Marketing wants a new dashboard. Insights wants a new cut of the data. A regional team wants a new market added. Access control and flexibility leads to scope creep and complexity. Someone has to triage, prioritise, build, test, and ship – forever. Vendors fund that roadmap across an entire client base; you fund it alone, in perpetual competition with every other engineering priority in the company.
16. Are you accounting for the fixed costs vendors amortise?
Security audits. GDPR compliance. Penetration testing. Privacy reviews. Uptime monitoring and incident response. Documentation and training. A vendor spreads these across all clients; an internal build carries every dollar directly. These are precisely the costs that don’t appear in the demo-stage business case and reliably appear in year two.
“Those maintenance costs are spread across the vendor’s entire client base, not sitting entirely on your P&L.”
Vijay Raj, Ex-EVP of CMI, Unilever
Buy vs. Build: Strategic Decision-Making for Future-Ready Organisations webinar
The expertise and scale problem. Knowing what good looks like (Questions 17-19)

17. Where does your cross-brand, cross-category expertise come from?
The hardest part of consumer insights AI isn’t the AI – it’s knowing what good looks like. What’s a meaningful shift in brand equity versus noise? Which review themes actually predict sales impact? How do you separate a lasting trend from a fad? That judgment is built by working across many brands, categories, and markets. i-Genie.ai was founded by veterans who led consumer insights at Unilever and Coca-Cola, and our platform encodes lessons from working with brands like Bayer, Danone, L’Oréal, Kenvue, Haleon, Proximo and Coca-Cola. An internal team learns only from its own company’s data – a sample size of one.
18. Can you replicate global coverage market by market?
Consumer signals are radically local. E-commerce leadership differs in every country – Amazon in the US, Flipkart in India, Mercado Libre in Latin America, Coupang in Korea. China is an ecosystem unto itself: Tmall, JD, Douyin, Xiaohongshu, and WeChat, with essentially no overlap with Western platforms. Each market means different data sources, different integrations, different languages, and different models tuned for accuracy in each one. We monitor insights in over 30 countries, with finely tuned models for over 20 languages. Building that coverage internally isn’t one project – it’s fifty.
19. Is this capability actually your competitive moat – or is acting on insights the moat?
The strongest argument for building is differentiation. But be precise about where differentiation lives. Collecting and processing digital signals is increasingly commoditised infrastructure; what’s differentiating is your proprietary first-party data, your decisions, and your speed of action. Salesforce’s guidance to enterprises applies here: buy the commoditised layer and reserve internal resources for what’s genuinely unique to you.
“Only do what only you can do.”
Vijay Raj, Ex-EVP of CMI, Unilever
Buy vs. Build: Strategic Decision-Making for Future-Ready Organisations webinar
Owning a data pipeline doesn’t win share. Acting on insights faster than competitors does. Building on a platform that turns signals into decisions through proven interfaces and highly tuned agentic tools like Presto is a faster way to win.
The honest question. The math your CFO would run (Question 20)

20. If your internal build has a 1-in-3 success rate, what’s the expected cost of being wrong?
Run the expected-value maths your CFO would run. MIT’s research puts internal AI builds at roughly a 33% success rate versus 67% for specialised external solutions. Multiply your fully loaded build cost by the probability of failure, add 18-24 months of opportunity cost while competitors act on insights you’re still building infrastructure for, and compare that to a vendor contract you can deploy in weeks and exit if it underperforms.
For most consumer brands, the answer isn’t ideological – it’s arithmetic.
Build vs. Buy Consumer Insights AI: Not ideological, arithmetic

None of this means “never build.” Some organisations, with genuinely differentiated data, highly talented engineers and data scientists, and sustained executive commitment, should build specific components – often on top of a bought foundation.
“Building internally is often the only way to own what you create.”
James Cummings, Ex-SVP of Global Head of Consumer & Business Intelligence, Kenvue.
Buy vs. Build: Strategic Decision-Making for Future-Ready Organisations webinar
The right answer is usually a hybrid, and the wrong answer is deciding by instinct rather than analysis.
That’s exactly the framework we walk through in our webinar, including real cost models, the conditions where hybrid approaches win, and how to make the case to your CFO and CTO whichever path you choose. For a deeper written treatment, see our companion guide: Should You Buy or Build AI for Consumer Insights?.
And if you’d rather spend the next 90 days generating insights instead of building pipelines – or, worse, caught in requirements gathering before even beginning – talk to us. We’ve already asked ourselves all 20 questions, every day, for years, on behalf of the world’s leading consumer brands.





























































