
Week five. That’s when it happens. Not in the model. Not in the demo. The timeline blows up in week five, almost every time, because the data isn’t ready — and nobody budgeted for making it ready.
This isn’t a theory. Teams that underestimate data engineering consistently blow their timelines in week five. The reason is straightforward: when data is unstructured or unlabeled, 30–40% of total project budget goes to data engineering before any model work begins. That figure rarely appears in the initial quote.
If you’re evaluating a custom AI engagement right now, this single fact should change how you read every proposal on your desk.
Listen To The Podcast Now!
The Cost Layer Nobody Quotes:
Ask ten AI vendors for a project estimate. Nine of them will describe model selection, training compute, and deployment infrastructure. What they won’t itemize — or will bury in assumptions — is the work that must happen before any of that. Cleaning and labeling raw data. Mapping source schemas to output schemas. Deciding what “good” even means for a model that has never seen your specific data before.
McKinsey’s 2025 State of AI report notes that AI high performers are more likely to define when model outputs require human validation. That means they built evaluation infrastructure, not just a model. Evaluation infrastructure takes time and budget. Most fixed-price proposals don’t honestly absorb it before discovery.
The pattern repeats. A document-extraction POC works beautifully on a curated PDF set. It hits production. The vendor’s output schema was never mapped to the warehouse management system. The project timeline doubles. This specific failure mode isn’t a hypothetical — it’s what happens when integration architecture isn’t locked down before a single line of model code is written.
Why Week Five, Specifically
The first four weeks feel fine. Scoping calls happen. A proof-of-concept materializes. Stakeholders are impressed. Then the team tries to connect the model to the actual system it’s supposed to live in and discovers the data it trained on looks nothing like what production will deliver.
Legacy systems are the usual culprit. Oracle databases. Homegrown CRMs with no documentation. Each one adds integration scope that no fixed-price quote can honestly absorb before discovery. The team pivots to data cleaning. Weeks stretch. The budget line for “model development” quietly funds a different kind of work entirely.
This is also where the handoff problem compounds everything. Many AI shops staff projects with ML specialists who hand off to integration engineers after model training. The seam between those two roles is where scope leaks. The ML team built something that works on its inputs. The integration team inherits a system never designed around the production data pipeline. Neither team is wrong. The organizational structure is.
The Evaluation Problem Nobody Talks About:
Even when integration goes smoothly, most teams have no way to measure whether the model is working in production. Not “working” in the demo sense. Actually catching every failure mode it was supposed to catch.
Consider the spam false-positive problem. A well-intentioned inbox classifier quietly routes real messages to spam. Globussoftai’s inbox intelligence agent catches 100+ misclassified messages per day — a failure mode only discoverable because an evaluation harness was built to measure it. Without that harness, the model looks fine in staging and breaks silently in production.
The benchmark that matters: build 100–500 real examples to measure accuracy and hallucination rate on every change. Most teams skip this. It’s not glamorous. It doesn’t show up in a demo. Skipping it means a user finds out the model is wrong before the engineer does.
The numbers behind this failure are stark. MIT Project NANDA’s July 2025 findings show 95% of organizations deploying generative AI fail to reach production or deliver ROI. Only 6% of organizations, per McKinsey, capture meaningful enterprise-wide value from AI. The data engineering tax and the missing evaluation layer together explain a large share of that gap.
Three Model Paths — and Where Each Hides Its Cost:
In 2026, three dominant paths exist. You can prompt-engineer a foundation model — fastest to ship. Fine-tune an open-weight model — best for narrow, repeatable tasks. Or build a RAG layer. Each hides its cost somewhere different.
Foundation model prompting is cheapest to start and most deceptive. The prompt works in development. In production, edge cases accumulate. You end up building an evaluation harness anyway — just later, under pressure, after something has already broken.
Fine-tuning demands labeled data. Which means data engineering. Which means week five.
RAG gets misunderstood most consistently. Teams spend weeks debating foundation model selection when the real accuracy driver is elsewhere entirely. For RAG systems, chunking strategy and embedding quality drive more accuracy gains than the choice of foundation model. A better embedding model on a mediocre chunking strategy will underperform a decent embedding model on a well-designed one. Most proposals don’t mention chunking at all — which tells you everything about how much the vendor has actually shipped.
Pilots vs. Full Builds: The Sizing Question:
The week-five blow-up often happens because the team committed to a full build when a pilot would have surfaced the data problems at a fraction of the cost. A focused pilot — one workflow, one integration — typically costs low tens of thousands and takes 6–10 weeks. A production platform touching multiple systems is a multi-quarter engagement. The difference between them isn’t ambition; it’s discovery.
Roughly one in three AI project ideas are killed in the discovery and feasibility phase, which takes one to three weeks. Killing a bad idea in week two costs three weeks. Killing it in week eighteen costs a quarter. The teams that avoid catastrophic timeline overruns aren’t luckier — they front-loaded the hard questions.
48% of companies already report positive ROI from AI investments, per a 2024 McKinsey report, and the distribution skews heavily toward teams that ran tight pilots before committing to full builds. That’s the pattern, not the exception.
What the Embedded Approach Actually Changes?
Handoffs create seams because they create information loss. When you hire a vendor to train the model and a separate team to integrate it, the integration engineers don’t know why the model was trained on certain data. The ML team doesn’t know what the production system actually emits. Nobody is lying. The structure is just wrong.
Globussoftai‘s embedded pod eliminates that boundary by design. One engineer drops into the client’s Slack and codebase, from scoping through production — no offshore handoffs, no relay between specialists. Integration architecture, including API auth to CRM, output schema, and escalation rules, gets signed off in weeks one and two. A first version deploys to real users by week three. Not staging. Real users.
This matters because data problems surface faster when the same person writing the model is also connecting it to the production pipeline. Week-five surprises become week-two conversations. The engagement is also built around multi-model flexibility — the team works across Claude, GPT, Gemini, and Qwen depending on which fits the task — rather than being locked to one provider’s stack.
The cost structure reflects what’s actually accessible. Frontier-lab FDE programs from Microsoft, Google Cloud, and Anthropic Solutions run $500k–$2M per year per embedded engineer, with contract minimums starting at $250k. The embedded pod from Globussoftai is priced at approximately one-tenth of that. For founders and mid-market teams who can’t qualify for those programs, it’s the only embedded AI engineering model that fits inside a real budget.
Four Questions to Ask Before You Sign Anything
These four questions will tell you more than any proposal document:
- What does your data engineering line item look like? If the vendor can’t answer this before discovery, you’re absorbing the timeline risk yourself.
- Who handles the integration to our production systems — and are they the same person who trains the model? A “yes” to the second half is rare and genuinely valuable.
- How do you measure whether the model is working in production, not just in staging? Ask specifically about evaluation set size and which failure modes it’s designed to catch.
- When does the first version hit real users? If the answer is “after staging sign-off” with no defined timeline for that sign-off, you’re looking at an open-ended engagement.
The data engineering tax is real. It’s predictable. The teams that escape week five are the ones who priced it honestly in week one — and structured the engagement so the same people who built the model are responsible for making it work in production.
If you’re exploring what a custom AI agent actually costs to build and deploy, the discovery phase is where the data question gets answered — and where bad projects get killed cheaply instead of expensively.
Ready to find out what your AI project really costs before week five finds out for you? Talk to the Globussoftai embedded team and get your scoping started.







