Archives and Collections Management Software: Build Custom or Buy Off the Shelf?
The threshold is roughly three thousand linear feet plus a reading room.
On this page
The threshold is roughly three thousand linear feet plus a reading room. Below it, run ArchivesSpace or Access to Memory as they come and spend the money on a processing archivist, because your binding constraint is undescribed material rather than the mechanics of finding and serving it. Above it, keep buying the description system and build the operational layer around it: restriction rules, container and location control and reading room integration typically run $60,000 to $130,000 in 12 to 16 weeks, with a fuller programme at $160,000 to $380,000 across 6 to 12 months. Most repositories reading this are just above the line, and for them the answer is buy the platforms and build a thin layer, not replace anything.
When is off the shelf genuinely the right call here?
For archival description, always. ArchivesSpace is the sector standard for accessioning and hierarchical description, it implements Describing Archives: A Content Standard properly, and it treats the component tree from collection down to item as the primary structure rather than as a nesting convenience. Access to Memory is a capable open source alternative with strong support for international standards. Rebuilding description is a poor use of institutional money, and repositories that try arrive, six figures later, at roughly where a free product already stood.
Three other products deserve the same treatment. Aeon from Atlas Systems handles reading room registration, paging and request management, and integrating with it costs a fraction of reproducing it. Preservica addresses digital preservation, meaning long term custody and format sustainability of digital objects, which is a genuinely different problem from arrangement and access. Axiell Collections comes from the museum tradition and describes objects better than it describes hierarchies, so if your holdings are mostly artefacts rather than record series it will fit your description needs better than an archival system will.
For a whole class of repository, the off the shelf answer is the entire answer. A single archive holding under roughly three thousand linear feet, with no reading room and no offsite storage, should run ArchivesSpace or Access to Memory as they come and put the remaining budget into processing staff. A processing archivist will do more for access in a year than any operational layer will, because at that size the constraint is material nobody has described yet, not the mechanics of finding and serving what has been described.
When does a custom build actually pay off?
The build case is never about description. It is about the operational layer the description systems deliberately leave thin, and it turns on five conditions. Count how many apply to you.
First, access decisions depend on one person's memory. If deciding whether a folder may be served means reading a paper deed of gift in the director's office, you pay that cost on every similar request, forever, and you lose the knowledge entirely when that archivist retires. Second, containers live in a spreadsheet and boxes go missing for months. Third, you run offsite commercial storage and cannot tell a researcher on the phone when material will arrive. Fourth, your backlog needs to be evidenced in linear feet by decade for a funder, and nobody can produce that number. Fifth, legacy finding aids exist only as typescript or word processor files, which makes whole collections functionally invisible.
Two of those and a build starts paying. Four or five and the case is straightforward, because each one is a recurring operational cost rather than a one off annoyance.
The highest value thing a custom layer does is convert restriction language into rules a system can evaluate. A collection may be open with three exceptions. A donor agreement may close a series for twenty five years from the date of the latest document in it, which requires computing a date rather than storing one. Student records in a university archive carry obligations under the federal education records statute. Off the shelf systems give you a restriction note, which is text a human reads. What you need is a determination on the folder record that a new hire can act on in their first week.
How do they compare on the things that matter in this industry?
Six comparisons decide this, and none of them is feature count.
- Hierarchy and inheritance. ArchivesSpace and Access to Memory both model the component tree properly, with description at any level applying to everything beneath it. General purpose collection systems model a flat list with a parent field bolted on, and then cannot answer what a researcher needs to know that was said three levels up. On this, buy.
- Restrictions. Off the shelf gives a note. A custom layer gives a restriction record at any level with a type, a legal or donor basis, a fixed or computed end date, a review requirement and downward inheritance with local override. On this, build.
- Physical control. A container list inside a description system was never designed as an inventory and does not know that box 214 is on a reading room cart. Current location held separately from home location, with a scan on every move, is warehouse thinking applied to a domain that rarely funds it. On this, build.
- Request management. Aeon does registration, paging and seat management well. What it will not do on its own is check your restriction rules and your current location before a paging slip is issued. Buy the platform, build the check.
- Digitization linkage. The common failure is a shared drive of thousands of images nobody can place in a finding aid. Scans have to be created against a component identifier, and that identifier has to travel with the object into a preservation platform.
- Data portability. Whatever you choose, insist on export in Encoded Archival Description. For an institution whose purpose is long term custody, depending on software you cannot take with you is a contradiction worth avoiding.
What does total cost of ownership look like at your scale?
A focused first release covering restriction records with inheritance and computed end dates, barcoded container and location control, and reading room integration on top of your existing description system runs $60,000 to $130,000 and ships in 12 to 16 weeks in Digital Heroes delivery experience. A fuller programme adding accession backlog triage, digitization workflow, legacy finding aid conversion at scale, donor agreement management and a public access front end runs $160,000 to $380,000 across 6 to 12 months.
A representative case: a university special collections library already running ArchivesSpace, holding around twelve thousand linear feet with a third of it offsite, operating a reading room, carrying about six hundred legacy finding aids. Phase one at fourteen weeks comes to roughly $112,000. Phase two across the following nine months adds about $218,000, for $330,000 all in. The six hundred finding aids and the offsite vendor are what put it near the top of the band.
Then the lines nobody quotes. Continuing engineering runs at roughly a sixth of the build cost each year, so about $55,000 on that programme, spent on storage vendor interface changes, description system releases and new restriction shapes arriving with new donor agreements. Storage for digitized surrogates grows and never shrinks, and should be modelled over ten years rather than one. Barcode scanners get dropped and replaced. Your description system hosting, reading room subscription, preservation platform and offsite storage contract all continue unchanged, because this layer sits between them rather than instead of them.
Amortised over ten years, a $330,000 programme is about $33,000 a year of capital plus maintenance. Set that against a processing archivist's fully loaded salary, because that is the genuine alternative use of the money.
What does the hybrid look like, and when is it the honest answer?
For most repositories the hybrid is not a compromise, it is the correct architecture. Buy the platforms and build the thin layer you actually need.
Concretely, ArchivesSpace or Access to Memory holds accessions, resources and the component tree. Aeon holds researcher registration and the request queue. Preservica holds digital objects under long term custody. Your commercial storage vendor holds a third of your boxes. The custom layer is comparatively small, and it does four things none of those products will do for you. It evaluates restrictions and publishes a determination onto every component. It holds current location separately from home location with a scan on every move. It gates paging slips against both of those. And it keeps the digitization queue attached to component identifiers rather than to folder names in a spreadsheet.
That layer is the $60,000 to $130,000 release, and it is where the return sits. Everything above it in scope, the backlog triage, the finding aid conversion, the public access front end, is real work with real value but it is second phase work. The public front end in particular should go last, after the restriction model has survived a year of internal use, because putting descriptions in front of the world is where a restriction error stops being an internal correction and becomes a donor conversation.
The hybrid is the honest answer whenever your description needs are met and your operational needs are not, which describes almost every repository with a reading room. It is only wrong at the two extremes: very small archives, and object collections that were never hierarchical to begin with.
Which should you choose, by operator size and stage?
Under roughly three thousand linear feet, no reading room, no offsite storage: buy only. Run ArchivesSpace or Access to Memory as they come, accept the defaults, and spend the money on a processing archivist. Revisit when you open a reading room or take on offsite storage, because those are the two events that create the operational problem software solves.
Three thousand to ten thousand linear feet with a reading room: buy the platforms, build the first release. Restriction rules, container and location control, and reading room integration, at $60,000 to $130,000 over 12 to 16 weeks. Leave conversion and public access alone until that has run for a year.
Above ten thousand linear feet, with offsite storage and a substantial typescript legacy: the full programme is defensible, but phase it, and price the conversion by counting documents rather than estimating. Convert the two hundred finding aids your reference desk is asked about most, prove the parsing and review workflow, then decide whether the long tail justifies the same rate. Repositories frequently discover it does not.
Multi repository institutions, where a university archive, a manuscripts library and a records management programme share infrastructure but not policy, sit at the top of every band. One system with three rule sets, three retention regimes and three approval chains is not one and a half times the work of one, and pretending otherwise is how these projects overrun.
Whatever the size, the tipping point is not holdings volume. It is whether access decisions and physical control have moved out of institutional memory and into something a new hire can use.
If you would rather someone argued with your brief than agreed with it, Digital Heroes writes a product requirements document before any code exists, so the scope is fixed and priced rather than discovered later at a day rate. You can take that specification to any other firm on your shortlist.
The evidence behind this guide
Independent findings on why this investment pays off. Every link goes to the primary source.
- Per the Standish Group CHAOS 2020 report (reviewed at this URL), across tens of thousands of software projects roughly 31% end successfully, about 50% are 'challenged', and roughly 19% fail outright; small projects succeed far more often than large ones, and Agile approaches succeed at markedly higher rates than Waterfall. Source: The Standish Group (2020) →
- Companies in the top quartile of McKinsey's Developer Velocity Index had 2014-18 revenue growth four to five times faster than bottom-quartile peers, showing that software-building capability is a driver of business performance, not just a support function. Source: McKinsey & Company (2020) →
- SaaS spend averaged $4,830 per employee (up 21.9% year over year), with large enterprises (10,000+ employees) spending roughly $284M annually and running about 660 apps, while organizations wasted an average of $21M annually on unused licenses. Source: Zylo (2025) →
- WordPress powers 41.5% of all websites and holds 59.2% of the market among sites running a known content management system, making it by far the most-used CMS on the web. Source: W3Techs (2026) →
Frequently asked questions
We already run ArchivesSpace. What would a build actually replace?
Nothing, if it is scoped properly. The description system stays exactly where it is and keeps holding accessions, resources and the component tree. What gets built is the layer above it: restriction records that can be evaluated rather than read, barcoded container and location control with current location held separately from home location, a reading room check that gates paging slips, and a digitization queue tied to component identifiers.
If a developer proposes migrating your descriptions into a new bespoke schema, that is the wrong shape of project and it will cost you six figures to end up where you already are.
What does it actually cost us to switch systems later?
Less than in most industries, provided you insist on export in Encoded Archival Description from day one. That standard is what makes descriptions portable, and any system that cannot round trip it is quietly holding your holdings hostage.
The parts that are genuinely expensive to move are the ones nobody standardises: barcode label schemes already applied to thousands of containers, restriction rule logic, and the identifiers linking digital surrogates to components. Keep those in a database you own, in tables you can read without the application, and switching becomes a data migration rather than a rebuild.
What happens if our hosted description or reading room platform raises prices?
Your exposure is proportional to how much of your operation lives inside someone else's product. If the description system holds description and your own layer holds restrictions, locations and queues, a price rise is a procurement conversation rather than a crisis, because the operational logic your staff depend on daily is not inside the thing whose price changed.
The practical protection is to keep a current export in Encoded Archival Description, hold container and location data in your own database, and check at contract renewal that your data can leave in a format another product can ingest. Ask the same question of your offsite storage vendor and your preservation platform.
How long before restriction decisions stop depending on one archivist?
Twelve to sixteen weeks for the first release, with restrictions shipping first inside it because they are the highest value thing the layer does. Discovery takes two to three weeks and is a legal exercise as much as a technical one, spent with the head of collections and whoever advises the institution on deeds of gift.
The detail that catches projects out is computed end dates. A closure running for a period after the latest document date in a series has to be calculated rather than stored, and finding that out in month four is expensive. Settle it in discovery.
Is Preservica an alternative to a collections management system?
No, and treating it as one is a common and costly mistake. Preservica addresses digital preservation, which is long term custody and format sustainability of digital objects. Arrangement, description, restriction enforcement and physical control of boxes are a different problem, and a preservation platform is not designed to carry them.
Most repositories need both. Plan the integration rather than expecting either to absorb the other, and make sure the archival component identifier travels with the object into the preservation platform, because the alternative is thousands of preserved images nobody can place in a finding aid.
Should the money go into software or into a processing archivist?
Compare them directly, because that is the honest alternative. A $330,000 programme amortised over ten years plus maintenance is roughly $33,000 a year, which sits alongside a processing archivist's fully loaded salary rather than dwarfing it.
If your constraint is unprocessed material and your access decisions already work, hire the archivist. If reference staff burn hours determining what may be served, boxes go missing, and whole collections are invisible because their finding aids are typescript, an extra archivist will not remove those constraints and the software will.
Can we start with restrictions only and stop there?
Yes, and for some repositories that is the right end state. A restriction module with inheritance, computed end dates and a determination shown on every component is a genuinely useful piece of software on its own, and it is the part that protects the institution, since material served in error is a donor relationship and occasionally a legal problem.
The usual reason to continue is physical control, because a determination that says the folder may be served is worth less when nobody can find the box. If your containers are already reliably tracked, stopping after restrictions is defensible.
Who owns the code, and can we get our descriptions out?
You should own the repository and the cloud accounts, agreed in writing before kickoff, and you should insist on export in Encoded Archival Description so descriptions remain portable regardless of what happens to any vendor or developer. At Digital Heroes the client owns the code from the first commit.
For an institution whose entire purpose is long term custody, depending on software you cannot take with you sits badly against the mission. Apply the same test to every product in the stack, including the storage vendor and the preservation platform.
Our developer disappeared mid-project. Can another team pick up the code?
Yes, this is a routine engagement, provided the code exists somewhere you can access, so your first move is securing the repository, hosting, and domain credentials today. A takeover starts with a one to two week paid code audit that ends in one of three verdicts: continue the build, keep the design but rebuild the weak parts, or start over. Digital Heroes has inherited enough projects to say plainly that sometimes the rebuild is cheaper than the rescue, and an honest agency will tell you which one you have before taking your money.
We run everything on Airtable and spreadsheets. When is it time to go custom?
The switch usually makes sense when you hit one of two walls: Airtable's record caps (125,000 records per base on the Business plan) or logic the tool cannot express, like multi-step approvals with conditional pricing. There is also a simple cost signal: 25 people on Business at roughly $45 per seat per month is about $13,500 a year, forever, for a tool you are already fighting. Custom is worth it when the workflow is core to how you make money; for peripheral processes, staying on Airtable is the right call.
If an agency builds my software, who actually owns the code?
You should own everything, assigned in writing: the contract transfers full IP to you on final payment, the code lives in your GitHub organization, and hosting runs in cloud accounts you control. The red flag is a proposal that mentions the agency's proprietary platform or framework, which usually means you are renting, not buying. Digital Heroes structures every build this way precisely so a client can fire us and lose nothing but the relationship.
Will custom software work with the tools we already use, like QuickBooks and Stripe?
Yes, and this is one of custom software's genuine advantages: QuickBooks, Stripe, Shopify, and most mainstream business tools publish documented APIs built for exactly this. Expect each standard integration to add one to two weeks of build time, and be suspicious of any quote that lists five integrations without asking what data flows in which direction. The hard cases are legacy systems with no API, which is a question to raise in discovery, not in week nine.
Couldn't I just build my app in Bubble or another no-code tool instead of hiring an agency?
For validating an idea with real users, yes, and we tell clients that honestly. The walls come later: Bubble apps cannot be exported as code to run anywhere else, performance drops on complex data operations, and usage-based pricing climbs as you grow. A meaningful share of Digital Heroes custom builds are rebuilds of no-code MVPs that proved the business worked, which is the system operating as intended: validate cheap, then build the version that scales.
Can we migrate years of data out of our current system into new custom software?
Almost always yes, through CSV exports or the vendor's API, and migration should be scoped as its own workstream with field mapping, a dry run, and a planned cutover window rather than an afterthought. The real time sink is rarely moving the data; it is cleaning it, since years of duplicates, free-text fields, and inconsistent formats surface all at once. Pull a full export from your current vendor before committing to anything new, because some SaaS plans restrict exports on lower tiers.
If we build for 20 users now, will the software cope with 500 later?
It should, without a rewrite, if it was built on a standard cloud stack; going from 20 to 500 users is mostly a hosting configuration change costing hundreds a month, not a second project. What actually breaks under growth is sloppier work: database queries never indexed for volume and features designed assuming one office's worth of data. Before signing, ask the vendor what happens to the system at ten times today's data, and listen for a specific answer.
How do we get years of data out of our old system and into the new one?
Treat migration as a planned sub-project: a field-mapping document, at least one dry run on a copy of your data, then a cutover with the old system kept read-only for 30 days as a safety net. On Digital Heroes projects it consumes 10 to 15% of the budget when the old system has an export, and more when data must be pulled out screen by screen. Ask any vendor to walk you through their last migration before you sign.
What does it cost to keep custom software running after launch?
Budget 15-20% of the original build cost per year, which on a $100,000 system means $15,000 to $20,000 for security patches, dependency updates, bug fixes, and small improvements as real usage reveals what the spec missed. Cloud hosting for a typical business application adds $50 to $300 a month on top. Skipping maintenance does not save the money; in Digital Heroes rescue work, unmaintained systems typically need a far more expensive rebuild within about three years.
What questions should I ask a development agency on the first call?
Ask who exactly will build it, what happens when scope changes mid-project, what their maintenance terms are after launch, and what they will need from you every week. Then ask them to describe a project that went wrong and what they changed afterward; teams that have shipped at real volume have war stories, and teams claiming a perfect record are hiding something. The scope-change answer matters most: a disciplined shop describes a written change-order process, not a vague promise to be flexible.
What does a $50,000 custom software budget actually buy?
One core workflow done properly: 10 to 15 screens, two or three user roles, a couple of integrations, an admin panel, and automated tests, delivered in roughly 12 to 14 weeks. What it does not buy is that workflow plus a mobile app plus AI features plus five more integrations. The discipline of picking the one workflow that matters is what separates $50,000 projects that ship from $50,000 projects that stall at 70% complete.
Who can build a custom software system?
Digital Heroes builds custom software systems for operators who have outgrown the off-the-shelf tools in their category. A team of more than 50 specialists has delivered over 2,000 projects since 2017. Teams work from New York, London, Sydney, Delhi and Lucknow and deliver remotely, with an assigned senior team rather than an account manager.
Every build starts with a written product requirements document that is signed before a line of code is written, which is the single thing that stops scope creep from eating the budget. Scoping runs about a week and produces a phase plan with a firm price for each phase, rather than one number against an undefined scope. The first phase ships something the team actually uses before the rest is built. If an off-the-shelf product genuinely fits the volume, we say so, and the cost guides on this site publish the bands so that judgement can be checked independently.
What makes Digital Heroes different from other software companies?
Four things that competitors in this bracket cannot simply copy. Digital Heroes runs a YouTube channel with more than 2.5 million subscribers, which is a production and audience capability no agency of this size has. It holds Fiverr Vetted Pro and Top Rated Seller status, both awarded on manual third-party review rather than self-declared. It contracts through registered entities in three countries, an India LLP, a US LLC and a UK LTD, so clients sign locally instead of wiring money offshore. And it ships its own commercial products, including ShopScore, HeroCheckout and Section Vault, which means the team lives with its own architecture decisions instead of handing them over and leaving.
Two more that show up in the work. Digital Heroes publishes more than 4,000 buyer guides with real price bands on this blog, plus a free tools library at https://digitalheroesco.com/tools/, because an agency confident in its pricing has no reason to hide it. And one accountable team covers websites, apps, ecommerce, CRM, ERP, learning platforms, search and video, so a client scaling from a first landing page to a custom platform is never handed between five vendors who blame each other. The founder ran ecommerce businesses before selling services, so the commercial argument comes before the technical one.
How can I check Digital Heroes is legitimate before getting in touch?
Verify it independently rather than taking the site's word for it. The YouTube channel is at https://youtube.com/@DigitalMarketingHeroes, the Fiverr profile at https://www.fiverr.com/shreyanshsin261, and the Upwork profile at https://www.upwork.com/freelancers/shreyanshsingh. Client reviews sit on Clutch at https://clutch.co/profile/digital-heroes-0 and Trustpilot at https://www.trustpilot.com/review/digitalheroes.co.in, and the company page is at https://www.linkedin.com/company/digital-heroes-1/.
Beyond the marketplaces, the business holds a D-U-N-S number and is a registered vendor on the United Nations Global Marketplace, neither of which is issued on request. Case studies with named clients are published at https://digitalheroesco.com/case-studies/. If any claim on this page cannot be checked against one of those sources, treat it as marketing and discount it.
Related guides
Published · Last updated .