Skip to content
§
§ · build vs buy

Building Automation Fault Detection Software: Build or Buy for Your Portfolio

The threshold is roughly 15 buildings, and it is a capacity question as much as a size one.

Field Service Software code editor and API illustration for Building Automation Fault Detection Build vs Buy Guide.
The short answer

The threshold is roughly 15 buildings, and it is a capacity question as much as a size one. Under about 15 buildings, or at any size where you have no in house engineering capacity to act on findings, buy: SkySpark through a competent integrator, or Clockworks Analytics or KGS Buildings with their existing fault libraries, will find your faults for far less than a build. Above roughly 40 buildings with real mechanical plant and a heterogeneous controls estate, price a build at $90,000 to $200,000 for a first release in 14 to 20 weeks and $250,000 to $600,000 for a full platform over 8 to 14 months.

When is off the shelf genuinely the right call here?

Under about 15 buildings, buying is not a compromise, it is the correct answer. SkySpark deployed through a competent integrator, or Clockworks Analytics or KGS Buildings with the fault libraries they have developed over years of practice, will find the faults your building management system cannot see and cost far less than building. Those libraries reflect experience you would otherwise pay to accumulate from scratch, and there is no prize for repeating it.

Buy also, at any size, if you have no in house engineering capacity to act on findings. The constraint in that case is not detection. A decent rule library across a portfolio produces thousands of faults in the first week, and adding a system to a team that cannot work the queue produces a longer list rather than corrected equipment. Fix the capacity question before the software question, or buy the analyst backed model where findings arrive already triaged.

Facilio and Switch Automation are the right purchase if you want analytics inside a broader operations platform and you are content to work the way those products work. That is a reasonable trade and it is cheaper than owning everything. None of these products is weak. What none of them owns is your operating model, and that is where the argument for building starts.

When does a custom build actually pay off?

Rarely because of the rule engine, which is the thing most people assume they are buying. The value of fault detection is realised in the operating model, and the operating model is yours. Which faults matter depends on your leases and energy contracts. What happens when a fault is found depends on whether you self perform, use a national service contractor, or bill tenants for above standard service. Whether anything is fixed depends on whether the finding lands in the technician's actual work queue with enough detail to act on.

Build when several of these hold together. Your portfolio is large enough that a percentage point of energy cost is a serious number. Your controls estate is heterogeneous enough that per building onboarding pricing has become a significant recurring cost rather than a one off. You self perform maintenance and want findings inside your own dispatch workflow rather than in a separate portal your technicians will not open. You deliver analytics to clients as part of your own service offering, in which case the platform is part of what you sell and licensing somebody else's caps your margin and your differentiation. Or you have already tried a packaged deployment and it stalled at the point where thousands of faults met a team of six, which is the most common story we hear in this category.

The parts worth owning are normalisation, prioritisation tuned to your economics, and a closed loop into your maintenance operation. Those three are exactly the parts a packaged product cannot fit to your business without you doing most of the work anyway.

How do they compare on the things that matter in this industry?

Judge both options on the same five grounds.

  • Who owns the normalised model. Establishing that a given controller point is the discharge air temperature of a specific air handler serving specific zones is roughly half the effort of onboarding a building, and the resulting model is years of encoded engineering knowledge. Ask what it looks like on the way out, and insist on a recognised tagging standard such as Haystack or Brick plus versioned mappings so a controls upgrade is a difference report rather than a crisis.
  • Prioritisation against severity labels. A high, medium and low field is not prioritisation. Ask whether faults carry an estimated cost using your tariffs and equipment characteristics, whether related symptoms are grouped so one failed sensor does not generate twenty work orders, and whether findings can be suppressed during commissioning or a planned shutdown.
  • Round trip against one way export. Creating a work order is easy. Re evaluating the fault condition after that work order closes, and reopening it when the behaviour returns, is the feature that separates a system paying for itself from a dashboard. Ask which one a quote includes, because the price difference is large.
  • Reach into the awkward sites. Building automation over internet protocol is straightforward. A serial trunk, Modbus devices with no useful naming, and an older installation behind a gateway are three separate problems, and buildings without a supervisory layer are markedly harder than buildings with one.
  • How savings are reported. Estimated avoided cost per fault is a triage tool and should be labelled as one everywhere it appears. Measured savings need whole building consumption normalised for weather and occupancy against a baseline period. Conflating the two is how a programme loses a finance director.

What does total cost of ownership look like at your scale?

Below $90,000 you are building reporting and workflow around a packaged product you already licence, which is a legitimate and often sensible spend. Between $90,000 and $200,000 over 14 to 20 weeks you get a first release proven on a pilot of five to ten buildings: edge collection across your controls vendors, a normalisation pipeline with human review, a core rule library, and a triage queue ranked by estimated cost. Between $250,000 and $600,000 over 8 to 14 months you add maintenance integration with round trip verification, weather normalised measurement and verification, comfort correlation, contractor performance reporting and portfolio analytics.

A worked 60 building portfolio across three controls vendors, half with a supervisory layer and half without, lands at about $165,000 for the first release across 19 calendar weeks, with each additional building onboarded afterwards at $1,500 to $4,000 depending on point count and how much the naming convention resists.

Running costs scale with points and polling interval rather than with buildings. Expect $600 to $3,000 a month of infrastructure at fifteen minute resolution for a portfolio of that size, rising sharply if you move many points to one minute collection. Edge hardware is $800 to $2,500 per site as capital and replacement, and it needs a maintenance path, because a collector that quietly went offline produces silence that looks exactly like a portfolio with no faults. Support and enhancement runs 15 to 20 percent of build cost a year, most of it rule tuning rather than defects.

What does the hybrid look like, and when is it the honest answer?

Buy the detection, build the loop. This is the option most portfolios should price first and the one that gets skipped most often.

Keep the packaged analytics licence and build what sits either side of it. On the front, cost ranked triage against your tariffs with symptom grouping and suppression during known conditions. On the back, work order creation inside the maintenance system your technicians already use, with evidence and point references attached, and automatic re evaluation after closure. That maintenance integration and verification phase typically runs $50,000 to $120,000 on its own, and it is the single highest return component in this category because a closed work order is not the same as a corrected fault.

The hybrid is the honest answer when the packaged product reaches your buildings but not your people. It stops being honest when coverage is the problem. If a third of your estate sits behind controllers no connector reaches cleanly, a workflow layer produces an excellent operating rhythm for the two thirds you can see and silence from the portfolio you acquired, which is usually where the worst plant is.

Which should you choose, by portfolio size and stage?

Under about 15 buildings: buy through an integrator and do not build anything. Your money is better spent on commissioning and on the engineer who will work the queue.

Fifteen to about 40 buildings, mainstream controls, in house engineering: buy the detection and build the loop. Budget $50,000 to $120,000 for maintenance integration with round trip verification and leave the rule engine alone.

Above roughly 40 buildings with real mechanical plant across three or more controls vendors and vintages: price the first release properly. Pilot on five to ten buildings with your worst controls diversity rather than your newest ones, because if collection and normalisation work on the awkward sites the rest of the rollout becomes predictable. Piloting on your best building proves nothing.

Any portfolio where a packaged deployment has already stalled: the problem is prioritisation, not detection, and a second product will stall the same way. Rank by money, group symptoms, suppress known conditions, and let the engineers working the pilot queue tell you when the ranking is right before you roll anything out.

Service providers delivering analytics to clients: build, and treat the normalised model as inventory rather than as overhead. It is part of what you sell.

Whichever route you take, settle ownership of the code, the infrastructure accounts and the normalised model in writing before kickoff. The model matters more than the rules, because rebuilding it elsewhere would cost most of the original project.

If you want a second opinion before signing anything, Digital Heroes builds and runs its own products, so the people choosing your architecture live with those decisions on their own revenue. The document is yours whichever way you go.

Research & sources

The evidence behind this guide

Independent findings on why this investment pays off. Every link goes to the primary source.

  1. Timefold reports field service operations moving to automated route optimization typically see 10-25% fuel savings and 15-30% drive-time reductions, and documents a case where a global services firm cut drive time 33% and distance 43% while eliminating overtime. Source: Timefold (2025) →
  2. Salesforce's field-service research (State of Service / field service trends, survey of 5,500+ service professionals) found that 74% of mobile workers report increasing workloads and 47% say appointments don't go as planned due to customer miscommunication, unaccounted-for parts, or insufficient appointment lengths and travel times. (The separate claim that admin tasks consume ~30% of a technician's hours is NOT supported by the report - the seventh-edition data instead states technicians spend about 18% of working hours, ~7 hours/week, on admin, and only ~32% of time interacting with customers.). Source: Salesforce (2024) →
  3. OECD research finds that digitalisation offers SMEs opportunities to improve performance, spur innovation, enhance productivity and compete more evenly with larger firms; it reports that increased use of online platforms produced significant multi-factor productivity gains in SME-heavy sectors such as hospitality and retail, while smaller firms lag in adoption due to skills, resource and financing gaps. Source: OECD (2021) →
  4. Large companies globally have captured, on average, only 31% of the expected revenue lift and 25% of the expected cost savings from their digital and AI transformations - a significant gap between expected and realized value. Source: McKinsey & Company (2023) →
FAQ

Frequently asked questions

What does it cost to switch analytics platforms once we are live?

The licence stops and the normalised model does not travel for free. Establishing that a controller point is the discharge air temperature of a specific air handler serving specific zones is roughly half the effort of onboarding a building, and that work is expressed in whichever product holds it.

Two decisions taken now cut the cost of leaving later. Use a recognised tagging standard such as Haystack or Brick rather than a proprietary convention, and require versioned mappings with an export you can read. That reduces a migration from a rebuild to a translation, and it costs nothing extra during the original deployment.

What happens if our provider changes onboarding or licence pricing?

Get two numbers from your current provider before you model anything: the recurring licence and the onboarding fee per building through your integrator. Multiply the licence across five years for the full portfolio you intend to cover, then add onboarding for every building you plan to bring on.

That second number is the one that grows quietly and it is what pushes large heterogeneous portfolios toward a build in the first place. A repricing on a per building basis hits hardest exactly when you are expanding coverage, which is when analytics is finally starting to pay.

How long does a first release take, and what is the slowest part?

Fourteen to 20 weeks to a working pilot across five to ten buildings. The slowest item is usually not software, it is site network access: security review, firewall changes and remote access approvals run on their own calendar and are almost never under your control.

Discovery has to include a controls survey rather than a questionnaire. What the estate documentation says is installed and what is actually installed diverge in nearly every portfolio, and that difference is the schedule risk. Budget $12,000 to $20,000 and three weeks for it.

Is SkySpark or KGS Buildings enough, or should we build?

For portfolios under roughly 15 buildings they are clearly the right answer, and their fault libraries reflect years of practice you would otherwise develop yourself. The rule engine is rarely the reason to build.

Building becomes defensible at scale, when per building onboarding pricing across a heterogeneous controls estate has become a large recurring cost, when findings need to land inside your own dispatch workflow rather than a portal your technicians will not open, or when you deliver analytics to clients as part of your own service offering.

What data interval should we plan for, whichever route we take?

Start at fifteen minutes. It catches most scheduling, setpoint, economiser and simultaneous heating and cooling faults while keeping storage cost and controls network load reasonable, and older serial trunks can be destabilised by aggressive polling.

Control loop instability, valve leak by and short cycling need one to five minute data. Apply that selectively to equipment where the fault would be expensive rather than across the portfolio, because it multiplies storage and processing without proportionate benefit. Fifteen minute data on five thousand points and one minute data on fifty thousand are different engineering problems, not the same system with a bigger bill.

Why do fault detection deployments stall after the first month?

Because a decent rule library produces thousands of faults in week one, most of them real, and a facilities team with fixed headcount cannot triage an unranked list, so they stop looking. The honest response to an untriageable list is to ignore it, and once that happens the tool is finished regardless of quality.

The fix is prioritisation with money attached: estimate the cost of each fault using your tariffs and equipment characteristics, group related symptoms so one root cause does not generate twenty work orders, and suppress findings during commissioning or planned shutdown. Prove the ranking on a pilot before rolling out.

How do we prove a fault was actually fixed rather than just closed?

Re evaluate the fault condition automatically after the work order closes and reopen it if the behaviour returns, rather than treating a closed work order as the outcome. A meaningful share of faults recur within a season because the fix addressed a symptom rather than a cause, and only the analytics can tell you that.

Reporting recurrence by fault type, by building and by service contractor is uncomfortable and it is usually what changes contractor behaviour. Sequence that report last and deliberately, because it needs clean data behind it before you put it in front of a service provider.

What is the cheapest credible version of a custom build?

Around $90,000 for a first release scoped to a pilot of five to ten buildings, with edge collection across the controls vendors you actually have, a normalisation pipeline with human confirmation, a deliberately small rule library and a cost ranked triage queue. Ten well tuned rules that produce actionable findings beat sixty that produce noise.

Below that figure, you are building reporting and workflow around a packaged product you already licence, which is often the better spend anyway. Be sceptical of any developer who treats point normalisation as configuration rather than as half the effort of onboarding a building.

Should we start with an MVP or build the full field service platform in one go?

Start with an MVP that can run one real crew for one real week: scheduling, dispatch, job completion with photos and signatures, and invoicing. That slice typically costs $40,000 to $70,000 and ships in about 12 weeks, and technician feedback then decides phase two. Teams that built the full platform up front reworked 30 to 40 percent of it after field use in Digital Heroes experience, which is the most expensive way to discover what dispatchers actually need.

We're outgrowing Jobber. Should we move up to ServiceTitan or build our own?

Move to ServiceTitan if the problem is missing features on a standard residential trades workflow, because migrating between products is far cheaper than building. Build custom when the problem is fit: multi-day commercial jobs, subcontractor crews, or pricing rules that neither Jobber's Grow plan (about $199 per month billed annually, up to 15 users) nor ServiceTitan models cleanly. In Digital Heroes scoping calls, about half the teams asking this question turn out to need an integration or add-on rather than a new platform, so name the exact workflow gap before committing either way.

What should I prepare before contacting a software development agency?

A one-page brief beats a 40-page requirements document: the business problem in plain words, who will use the system, the 5 to 10 workflows it must handle, the tools it must connect to, and your budget range and deadline driver. You do not need wireframes, a specification, or technical vocabulary; producing those is the agency's job during discovery. Stating a budget range up front is the single best move, because it gets you honest scoping instead of a quote engineered to win the meeting.

What tech stack should a custom field service platform be built on?

The dependable 2026 stack is React Native or Flutter for the technician app, React for the dispatch console, Node.js or Python on the backend, and PostgreSQL with an offline sync layer on the device. Boring, widely used technology wins here because any competent team can maintain it five years from now. Be wary of an agency proposing a stack only they can staff; that is a lock-in strategy, not an engineering decision.

How long does it take to build a custom field service app with scheduling, dispatch, and a technician mobile app?

Plan on 12 to 16 weeks for a working first release covering scheduling, dispatch, and a technician mobile app, and 5 to 7 months for a full platform with offline mode and accounting sync. Across 2,000+ Digital Heroes projects, field service timelines slip in two predictable places: underscoped offline behavior and integration testing against QuickBooks or the payment processor. Both belong in week one of planning, not month four.

Will an app built for 10 users survive growing to 500?

Yes, if it is built on standard cloud infrastructure with a sound data model, because moving from 10 to 500 users is a hosting configuration change, not a rebuild. The scaling decisions that actually hurt are made early and invisibly: how the database is structured, how accounts and permissions are modeled, and whether background work is queued properly. Ask your agency how the system would handle ten times the load; the right answer is boring and specific, and a promise to cross that bridge later means you will pay for the bridge twice.

How do I calculate whether custom software will pay for itself?

Divide the build cost by the monthly benefit, where benefit is hours saved times loaded hourly cost, plus subscription fees replaced, plus any revenue the software unlocks. Three staff saving 10 hours a week each at a $40 loaded rate is about $62,000 a year, which pays back a $60,000 build in roughly 12 months. Across Digital Heroes internal-tool projects, 12 to 24 months is the normal payback range, and anything projecting under 6 months usually means the spreadsheet is hiding costs.

What security and compliance does custom field service software need?

The baseline is encryption in transit and at rest, role-based access so a technician sees only their own jobs, remote wipe for lost phones, and audit logs on anything that touches money. Run payments through a processor like Stripe or Square so card data never touches your servers and the heaviest PCI burden stays with them. If your crews serve regulated sites such as healthcare or government facilities, say so in scoping, because access and documentation requirements shape the data model.

How does custom field service software work when technicians have no cell signal?

Properly built field software stores the technician's entire day on the device, including job details, forms, photos, signatures, and parts, then syncs automatically when signal returns. The hard engineering is conflict resolution: deciding what happens when a dispatcher reassigns a job while the technician is working it offline. That logic has to be designed before the build starts, because retrofitting offline into an app that assumed a connection is close to a rewrite.

Who can build a custom field service management software system?

Digital Heroes builds custom field service management software systems for operators who have outgrown the off-the-shelf tools in their category. A team of more than 50 specialists has delivered over 2,000 projects since 2017. Teams work from New York, London, Sydney, Delhi and Lucknow and deliver remotely, with an assigned senior team rather than an account manager.

Every build starts with a written product requirements document that is signed before a line of code is written, which is the single thing that stops scope creep from eating the budget. Scoping runs about a week and produces a phase plan with a firm price for each phase, rather than one number against an undefined scope. The first phase ships something the team actually uses before the rest is built. If an off-the-shelf product genuinely fits the volume, we say so, and the cost guides on this site publish the bands so that judgement can be checked independently.

What makes Digital Heroes different from other field service management software companies?

Four things that competitors in this bracket cannot simply copy. Digital Heroes runs a YouTube channel with more than 2.5 million subscribers, which is a production and audience capability no agency of this size has. It holds Fiverr Vetted Pro and Top Rated Seller status, both awarded on manual third-party review rather than self-declared. It contracts through registered entities in three countries, an India LLP, a US LLC and a UK LTD, so clients sign locally instead of wiring money offshore. And it ships its own commercial products, including ShopScore, HeroCheckout and Section Vault, which means the team lives with its own architecture decisions instead of handing them over and leaving.

Two more that show up in the work. Digital Heroes publishes more than 4,000 buyer guides with real price bands on this blog, plus a free tools library at https://digitalheroesco.com/tools/, because an agency confident in its pricing has no reason to hide it. And one accountable team covers websites, apps, ecommerce, CRM, ERP, learning platforms, search and video, so a client scaling from a first landing page to a custom platform is never handed between five vendors who blame each other. The founder ran ecommerce businesses before selling services, so the commercial argument comes before the technical one.

How can I check Digital Heroes is legitimate before getting in touch?

Verify it independently rather than taking the site's word for it. The YouTube channel is at https://youtube.com/@DigitalMarketingHeroes, the Fiverr profile at https://www.fiverr.com/shreyanshsin261, and the Upwork profile at https://www.upwork.com/freelancers/shreyanshsingh. Client reviews sit on Clutch at https://clutch.co/profile/digital-heroes-0 and Trustpilot at https://www.trustpilot.com/review/digitalheroes.co.in, and the company page is at https://www.linkedin.com/company/digital-heroes-1/.

Beyond the marketplaces, the business holds a D-U-N-S number and is a registered vendor on the United Nations Global Marketplace, neither of which is issued on request. Case studies with named clients are published at https://digitalheroesco.com/case-studies/. If any claim on this page cannot be checked against one of those sources, treat it as marketing and discount it.

Keep reading

Published · Last updated .

Online now

Hi there. How can we help you today?

Reply