Skip to content
§
§ · build vs buy

Reliability, FMEA and RCM Software: Build vs Buy

Buy. One site, a few hundred maintainable assets and one reliability engineer is a Hexagon Reliasoft or Isograph purchase, and the money left over belongs in condition monitoring hardware.

Internal Tools Development software overview illustration for Reliability Fmea RCM Software Build vs Buy Guide.
The short answer

Buy. One site, a few hundred maintainable assets and one reliability engineer is a Hexagon Reliasoft or Isograph purchase, and the money left over belongs in condition monitoring hardware. Build when studies keep ending as spreadsheets nobody keys into SAP Plant Maintenance or Maximo, because an approved task that never reaches a planner changes nothing on the floor.

What the off-the-shelf products actually do well

A reliability manager filters preventive tasks for the mill area and gets back four thousand. A third came from the original equipment manufacturer manuals at commissioning, a third were added after incidents by superintendents who have since left, and nobody can account for the rest. He has been asked to take twelve percent out of the budget. That is the problem, and for most plants the answer is still to buy a product rather than commission one.

Hexagon Reliasoft is the strongest thing available for Weibull life data analysis and formal reliability centred maintenance and failure mode and effects analysis (FMEA) worksheets under IEC 60812. Isograph Availability Workbench does reliability, availability and maintainability simulation properly, so if your question is production availability across a train of equipment it is the right tool and a build would be worse. ARMS Reliability OnePM is built around living strategy libraries with reuse across similar assets, which is the correct idea if your asset base is genuinely templated. GE Vernova APM is the broadest of the four. APIS IQ-FMEA is the automotive standard bearer.

These products also carry method and defensibility. SAE JA1011 sets the seven questions a process has to answer before it can be called reliability centred maintenance, and JA1012 explains them. A recognised tool producing recognised worksheets is worth something when a regulator, an insurer or a customer auditor asks how the interval was set.

So buy if you have one site, a few hundred maintainable assets and one engineer. Buy also if your functional locations are a mess, because software cannot fix a hierarchy where the same crusher appears three times under different codes. Fix the master data first with anybody, then talk about tooling.

Where they stop: the approved task never reaches a planner

Every one of those products stops in the same place, and it is the return trip.

A study is worthless until the approved task appears in front of a planner as a maintenance item, on a maintenance plan, with the right task list, work centre, materials and cycle. In SAP Plant Maintenance that means general task lists, maintenance items, maintenance plans, strategies and packages, each with master data rules your organisation customised years ago. In IBM Maximo it means job plans, preventive maintenance records and route stops. The specialist tools export a spreadsheet. Somebody then keys it in, and because keying four thousand rows is intolerable, the study stays in the folder and nothing changes on the floor.

The second thing they stop short of is your criticality framework. Criticality is not a number, it is a decision framework your business already owns, usually a risk matrix with your own consequence categories for safety, environment, production loss per hour and regulatory exposure. Every plant has customised it, and a generic criticality wizard produces an answer your risk committee will not sign.

The third is more specific and it catches suppliers to the automotive sector regularly. The 2019 AIAG and VDA FMEA handbook replaced the Risk Priority Number with Action Priority tables. If your tool still ranks by the multiplied score, your output does not match what a customer audit expects to see, and the argument happens in front of the customer rather than internally.

The fourth is the feedback loop. Take a failure mode you wrote a task for two years ago and try to show how many times it has occurred since, across that asset class, across all sites, without opening a spreadsheet. The evidence lives as free text in notification long descriptions, and nobody has time to read forty thousand of them.

The arithmetic: per seat licences versus a build at your asset count

These tools are licensed per named or concurrent seat, so the calculation is a straightforward one. Take the annual seat cost, call it S, and count who genuinely needs one. Reliability engineers, the maintenance engineers running studies at each site, and whoever maintains the templates. Call that N.

Then count the people who need the answer rather than the tool. Planners, superintendents, the safety department reviewing a task removal, the finance analyst holding the budget line. Under per seat licensing they either do without or you buy seats for readers, which is where the cost stops tracking the value.

The crossover sits at roughly 12 to 20 seats, or around 3,000 maintainable assets once more than one site is involved, whichever arrives first. Below that, buy the specialist tool and accept the manual write back as a known cost. Above it, template reuse across sites is worth real money and the write back gap becomes the binding constraint on whether any of the analysis reaches production.

Then price the thing nobody has in the model. Ask how many tasks in your current plan have no traceable parent failure mode. Take the labour hours attached to them for one year. That figure is not a saving you can promise, because the safety department will reinstate some of them, but it is the size of the question, and it is usually larger than every licence you hold.

What a custom build actually costs

Bands, from Digital Heroes delivery experience. A first release covering the failure mode library, your own criticality framework, task selection and interval logic, and validated write back into one maintenance system runs $70,000 to $150,000 and ships in 12 to 18 weeks. That is a system your reliability team uses against a real area, not a pilot. A full strategy platform adding work order text classification, Weibull fitting on real history, availability modelling, spares and bill of materials linkage, multi site template governance and a drift dashboard runs $180,000 to $450,000 over 8 to 14 months.

Data migration adds 10 to 25 percent and here it means importing existing studies and cleaning functional locations. If the same asset appears under three codes, that reconciliation is the migration and it belongs in the plan rather than in a hopeful assumption.

Year two runs 15 to 20 percent of build cost annually. Task lists change, a site is acquired, the risk matrix gets revised after an incident, and each is a maintenance ticket.

What pushes cost up here specifically: write back is the big one, because your maintenance plan master data has been customised and the transport and testing path through your landscape is not ours to control. Multiple maintenance systems after acquisitions, each with a different functional location convention, is second. Formal safety instrumented function work under IEC 61511, if you want protective device proof test intervals in the same system, is a specialist workstream and should be priced as one.

The four situations where building wins

  • Regulatory fit. If your intervals feed safety instrumented functions under IEC 61511, or risk based inspection under API 580 and API 581, or an ISO 55001 asset management system, the interval is an auditable commitment rather than a preference. Recording which method was used for each analysis, the approver and the date, and being able to show it, is the requirement. A worksheet in a folder does not meet it.
  • Scale economics. Seat counts rising with every site while template reuse across those sites is exactly what should be reducing the work. Paying more to do the same analysis again is the wrong direction.
  • A workflow that is your competitive advantage. Operators who genuinely run asset strategy as a discipline do it through their own criticality framework and their own equipment taxonomy, often aligned to ISO 14224. Standardising that to a vendor's wizard hands away the judgement your reliability function exists to apply.
  • Integration sprawl across three or more systems. A maintenance system, a historian, a condition monitoring platform, a spares and inventory module and a document store, each holding part of the case for an interval. Every pair is a manual export and the evidence question crosses all of them.

One of those is a vendor conversation. Two is a build.

How to decide in a week

Run a parentage audit on one area. Five days, no software purchase, and the result is a number your maintenance manager cannot dismiss.

Monday: export every preventive task for a single area from your maintenance system. Do not filter it. Count them.

Tuesday and Wednesday: for a random sample of fifty, try to find the parent. Which function does this protect, which functional failure, which failure mode, who set the interval and on what evidence, and when was it last reviewed. Record the answer as documented, remembered by someone, or unknown.

Thursday: take the tasks marked unknown and pull their labour hours and materials for the last twelve months. Then take the ten failure modes that caused the most unplanned downtime in that area and check whether any task addresses them at all. The gaps in both directions are the finding.

Friday: price it. Labour attached to unparented tasks, plus the downtime attached to unaddressed failure modes, against the licence line and the bands above. If most of the fifty were documented and traceable, your tool is working and you should buy more seats rather than a system. If more than half were remembered or unknown, the strategy layer does not exist and no purchase creates it.

What follows is a paid discovery phase rather than a proposal. Two to three weeks, fixed fee, ending in a signed product requirements document covering the data model from asset class through function, functional failure, failure mode, task and approval, the write back objects by name, and acceptance criteria. You own that specification whoever builds it.

Who we are wrong for: single site operators with a stable task list, anyone wanting Weibull analysis rebuilt, and anyone whose functional location hierarchy has not been cleaned. Digital Heroes writes that document before any code, with more than fifty specialists and India LLP, US LLC and UK LTD entities so intellectual property assigns under your own law. ShopScore, HeroCheckout and Section Vault are our own products, over 2,000 projects sit behind us, and you meet the named team before signing. We are listed on Clutch, Trustpilot, Fiverr Vetted Pro and D-U-N-S.

Book a 30-minute call with Digital Heroes and get a written plan and a fixed quote within 48 hours.

Research & sources

The evidence behind this guide

Independent findings on why this investment pays off. Every link goes to the primary source.

  1. Companies in the top quartile of McKinsey's Developer Velocity Index had 2014-18 revenue growth four to five times faster than bottom-quartile peers, showing that software-building capability is a driver of business performance, not just a support function. Source: McKinsey & Company (2020) →
  2. Salesforce research indicates sales reps spend only about 30% of their time actively selling, with much of the rest lost to administrative work including manual CRM data entry and updates. Source: Salesforce (2024) →
  3. Across more than 5,400 IT projects studied by McKinsey and the University of Oxford BT Centre, large IT projects ran on average 45% over budget and 7% over schedule while delivering 56% less value than predicted. Source: McKinsey & Company / University of Oxford (BT Centre for Major Programme Management) (2012) →
  4. IBM frames first-time fix rate as a core field service KPI, noting the industry average sits around 80% (roughly one in five jobs needs a return visit). Correction: IBM cites best-in-class providers at 89-98%, not '85%+'. Source: IBM (2024) →
FAQ

Frequently asked questions

How long before a reliability strategy build is useful?

Twelve to eighteen weeks for a first release when you scope it to one plant, one area and the equipment classes carrying the most unplanned downtime. The schedule risk is rarely engineering. It is master data, because a strategy layer sitting on a functional location hierarchy where the same crusher appears three times under different codes produces confident nonsense. Operations with a governed hierarchy move noticeably faster.

Who owns the code and the failure mode library if an agency builds this?

You should own the repository, the cloud accounts, the failure mode library and the right to hire any other firm, written into the contract before kickoff. At Digital Heroes the client owns everything from the first commit. In a discipline whose whole purpose is removing single points of failure, accepting a developer who holds your repository is a contradiction worth naming out loud.

What happens to our studies if we change maintenance systems later?

The strategy layer should survive it, which is an argument for separating the analysis model from the write back adapter. Failure modes, functions, intervals and approvals belong to you and are system independent. What changes is the mapping to job plans and preventive maintenance records, or to task lists, maintenance items and plans. Ask any developer to show that separation on a whiteboard before signing.

Can we keep Reliasoft and build only the write back layer?

Yes, and for teams already productive in a specialist tool it is the cheapest way to close the real gap. Keep the tool for worksheets and life data analysis, and build the adapter that turns approved tasks into the right maintenance objects with validation, batch error handling and a reconciliation report. The question to settle first is what happens when a validation fails halfway through eight hundred records.

What is the difference between FMEA and RCM in practice?

Failure mode and effects analysis catalogues how something fails and what the consequence is. Reliability centred maintenance uses that catalogue to decide what maintenance should exist, working from function through functional failure and failure mode to a task and an interval. Most operators run the full process on high criticality assets and a streamlined review elsewhere, and the system should record which method was used.

Should we score risk using RPN or Action Priority?

If you supply the automotive sector, Action Priority, because the 2019 AIAG and VDA handbook replaced the Risk Priority Number with Action Priority tables and customer audits expect the newer output. Outside automotive, use whichever your own risk matrix defines, but record the method against each analysis so an auditor can see the basis rather than inferring it from a number with no stated scale.

Can artificial intelligence classify our old work order history?

Yes, and it is the one place in reliability tooling where a language model genuinely pays. Notifications written as noisy drive end bearing, bearing knocking and non drive end bearing collapsed all map to the same failure mode, which lets you fit real intervals instead of assumed ones. Build a review queue where an engineer confirms or corrects, because the model will be confidently wrong on a share of records.

How do we stop the strategy drifting away from the maintenance system?

Version every analysis, require a management of change record to alter an approved task, and run a drift report comparing the approved strategy against what the maintenance system actually holds. People will edit plans directly under shutdown pressure and that is acceptable. What is not acceptable is nobody knowing, and a weekly drift report with a named owner turns a silent divergence into an agenda item.

Does removing preventive tasks create a safety exposure?

Only if the removal has no written justification, which is exactly why the parent chain matters. A task deleted because nobody could find its purpose is a risk. A task retired because the failure mode it addressed is documented, the consequence category is recorded and an engineer approved the decision with a date is a controlled change. The paperwork is what makes it survive contact with the safety department.

Is a spreadsheet ever acceptable for maintenance strategy?

For a single site with a few hundred assets, one engineer and a stable task list, yes, and disciplined spreadsheet governance beats a partly adopted custom tool. The line is crossed when template reuse across sites becomes the point, when the same analysis is being redone every few years because the last outputs became unfindable, or when nobody can trace what a given task prevents.

What are the biggest mistakes first-time software buyers make?

Choosing the lowest bid, paying more than 30-40% upfront instead of on milestones, skipping a written specification, and having no maintenance plan for after launch. The most expensive of the four in Digital Heroes rescue projects is the missing spec: without written acceptance criteria, done becomes an argument instead of a checklist, and every disagreement resolves in the vendor's favor. Fix those four and you have avoided most of the ways these projects fail.

Does it matter which tech stack the agency wants to use?

Yes, but not in the way most buyers expect: the goal is boring, popular technology such as React, Node.js or Python, and PostgreSQL, because any future team can maintain it and hiring a replacement developer takes days, not months. The red flag is an agency-proprietary framework or an unusual language, which welds you to that one vendor no matter what your contract says about code ownership. A useful test: could you find three freelancers fluent in this stack within a week? If not, push back.

What happens to my software if the agency shuts down or we stop working together?

Nothing dramatic, if the engagement was set up correctly: the code sits in your repository, hosting runs on your cloud account, and a handover document explains how to deploy and operate the system. Any competent replacement team can then take over in days rather than months. If the agency controls the repo, the servers, or the domain, fix that now, because renegotiating access during a dispute is the most expensive place to discover the problem.

How long does it take to build an internal tool from scratch?

A working first version typically ships in 4 to 8 weeks, and larger multi-module tools run 10 to 16 weeks. Across Digital Heroes internal tool projects the schedule splits into roughly one week of process mapping, 3 to 6 weeks of build, and 1 to 2 weeks of testing with your actual staff. The most common delay is not development but waiting on the client for sample data and workflow decisions, so name one internal owner before kickoff.

When does a company outgrow Airtable?

The usual breaking points are record limits, permissions, and automation complexity. Airtable's Team plan caps each base at 50,000 records and Business at 125,000, so operations logging thousands of rows a month hit the ceiling within a year or two. The other trigger Digital Heroes sees constantly is permissions: restricting who can view specific fields or records is clumsy below Airtable's Enterprise tier, which becomes a genuine problem once salaries, pricing, or client contracts live in the base.

Can we start on Airtable or Retool now and move to custom software later?

Yes, and it is often the smartest sequence: run the workflow on Airtable or Retool for 6 to 12 months to learn what you actually need, then go custom once the process stabilizes. The no-code version becomes free requirements documentation, and its data exports cleanly into a custom database. The one risk is waiting too long, because teams stack automations and workarounds until migration becomes a project of its own, so set a concrete trigger in advance, such as hitting Airtable's 50,000-record Team plan cap.

How many people should be working on my software project?

Three to five for a typical focused build: a project lead, one or two engineers, a designer, and part-time QA, which is the standard shape across 2,000+ Digital Heroes projects. Larger platforms justify 6 to 10, but a ten-person team on a small first version usually signals bill padding rather than horsepower. What predicts success is whether a senior engineer is writing your code daily, not the headcount on the proposal.

Who can build a custom internal tools system?

Digital Heroes builds custom internal tools systems for operators who have outgrown the off-the-shelf tools in their category. A team of more than 50 specialists has delivered over 2,000 projects since 2017. Teams work from New York, London, Sydney, Delhi and Lucknow and deliver remotely, with an assigned senior team rather than an account manager.

Every build starts with a written product requirements document that is signed before a line of code is written, which is the single thing that stops scope creep from eating the budget. Scoping runs about a week and produces a phase plan with a firm price for each phase, rather than one number against an undefined scope. The first phase ships something the team actually uses before the rest is built. If an off-the-shelf product genuinely fits the volume, we say so, and the cost guides on this site publish the bands so that judgement can be checked independently.

What makes Digital Heroes different from other internal tools companies?

Four things that competitors in this bracket cannot simply copy. Digital Heroes runs a YouTube channel with more than 2.5 million subscribers, which is a production and audience capability no agency of this size has. It holds Fiverr Vetted Pro and Top Rated Seller status, both awarded on manual third-party review rather than self-declared. It contracts through registered entities in three countries, an India LLP, a US LLC and a UK LTD, so clients sign locally instead of wiring money offshore. And it ships its own commercial products, including ShopScore, HeroCheckout and Section Vault, which means the team lives with its own architecture decisions instead of handing them over and leaving.

Two more that show up in the work. Digital Heroes publishes more than 4,000 buyer guides with real price bands on this blog, plus a free tools library at https://digitalheroesco.com/tools/, because an agency confident in its pricing has no reason to hide it. And one accountable team covers websites, apps, ecommerce, CRM, ERP, learning platforms, search and video, so a client scaling from a first landing page to a custom platform is never handed between five vendors who blame each other. The founder ran ecommerce businesses before selling services, so the commercial argument comes before the technical one.

How can I check Digital Heroes is legitimate before getting in touch?

Verify it independently rather than taking the site's word for it. The YouTube channel is at https://youtube.com/@DigitalMarketingHeroes, the Fiverr profile at https://www.fiverr.com/shreyanshsin261, and the Upwork profile at https://www.upwork.com/freelancers/shreyanshsingh. Client reviews sit on Clutch at https://clutch.co/profile/digital-heroes-0 and Trustpilot at https://www.trustpilot.com/review/digitalheroes.co.in, and the company page is at https://www.linkedin.com/company/digital-heroes-1/.

Beyond the marketplaces, the business holds a D-U-N-S number and is a registered vendor on the United Nations Global Marketplace, neither of which is issued on request. Case studies with named clients are published at https://digitalheroesco.com/case-studies/. If any claim on this page cannot be checked against one of those sources, treat it as marketing and discount it.

Keep reading

Published · Last updated .

Online now

Hi there. How can we help you today?

Reply