Setting up a college digital library: the software is free, the scanner is the only real hardware, and the cost nobody budgets is a person
🌐 इस लेख को हिन्दी में पढ़ें
In short: A step-by-step guide to setting up a digital library in an Indian college. It covers the four decisions that must be made before any quotation is sought, why Koha comes before scanning, how to switch on the free national resources first, standing up a DSpace repository with a usable metadata minimum, the deposit workflow that actually collects material instead of asking nicely, scanning specifications and storage, remote authentication, and licensing last against usage data — plus what the whole thing really costs.
Most college digital library projects begin with a quotation. A vendor demonstrates a portal, a figure is put in the budget, and eighteen months later there is a login nobody uses, a few hundred scanned pages nobody can find, and a librarian who has learned not to bring it up in meetings.
They begin at the wrong end. Almost everything that decides whether this works is settled before any software is chosen, and most of it is not technical at all. What follows is a sequence that survives contact with an actual Indian college — one where the librarian has three other jobs, the internet drops in the afternoon, and the principal wants to see something working this term.
Before any purchase: four decisions
Who owns this. One named person, with the time written into their duties. The single most reliable predictor of a stalled project is that responsibility was distributed across a committee. A committee can approve; it cannot switch anything on.
What you are allowed to publish. Decide, in writing, what will go into the repository and under what permission: theses and dissertations, question papers, syllabi, institutional journals, conference proceedings, annual reports, departmental photographs. Write the author permission form and the embargo rule before you accept the first deposit, because retrofitting consent onto material already online is unpleasant and sometimes impossible.
How a user will be recognised off campus. This determines whether anything you license is usable, so decide it before you license anything. It is a question for whoever runs the college's email and Wi-Fi accounts, and the answer you want is that one college identity will eventually open everything.
What "done" looks like. Pick two numbers you will report: items deposited, and sessions from off campus. Vague success is how a project quietly becomes nobody's problem.
Step 1 — Automate the catalogue
Koha, open-source and widely used across Indian institutions, is the standard choice. It costs nothing to license; you pay for hosting and, usually sensibly, for an implementation partner in the first year.
The work is not the installation. It is retro-conversion: getting the existing stock into machine-readable records. For a college of twenty to fifty thousand volumes this is months of careful data entry, and it is the stage projects underestimate by the widest margin. Two things make it tractable — copy cataloguing, where records are pulled from existing sources rather than typed from scratch, and accepting a minimum record for older stock instead of perfecting every entry.
Do this first even though it produces no digital content. Everything downstream hangs off a catalogue, and a college that scans before it catalogues ends up with an unsearchable heap.
Step 2 — Switch on what is already free
Before anything is bought, take what exists. This step costs almost nothing and is skipped constantly, because it is nobody's job.
Become an N-LIST member if you are a college — a modest annual fee for e-journals and e-books, and a login your students can use from anywhere. Check whether your institution falls under ONOS, which since January 2025 has funded journal access for government institutions and central R&D laboratories; if you are one, you may already have far more than you are using. Put NDLI, Shodhganga, Krishikosh and the other open resources into the library's own page, and — this is the part that actually moves usage — into the induction session every first-year student sits through.
Then register the college's IP range with every publisher you already have. Colleges routinely lose access they are paying for because the network changed and nobody told anybody. If your connection has no static IP, that is not a blocker; it is an argument for moving straight to a proxy at step 6.
Step 3 — Stand up the repository
DSpace is the common choice for an institutional repository and is also free. Hosted or on your own server both work; hosted is usually the right answer for a college with no systems staff.
Structure it the way the institution is actually organised — department, then material type, then year — because that is how people will browse it, and a structure copied from another college's screenshot will fight you for a decade.
Agree a metadata minimum and enforce it from the first item: title, author, supervisor or department, year, type, language, subject keywords, an abstract, and rights. Nine fields, filled consistently, will outperform thirty fields filled when someone remembers. Fix one form of each author's name and keep it; inconsistent name forms are the reason a repository cannot tell you what a department has published.
Ask for persistent identifiers rather than raw URLs. DOIs involve membership and a fee, which not every college will take on; handles are the usual free alternative. Either is better than a link that dies when the platform moves.
Step 4 — The deposit workflow, which is the whole project
Voluntary deposit does not work. This is the most consistent finding in the field and it is worth stating bluntly: a repository built on asking people nicely collects almost nothing after the first enthusiastic month.
What works is attaching deposit to a moment that already exists and already blocks something. The thesis submission and no-dues process is the obvious one — the electronic copy is deposited as part of clearance, not as a favour afterwards. Annual reporting, appraisal and accreditation documentation are the others: departments are already assembling lists of publications, and the repository can be where that list is produced rather than a second place to type it.
Two supporting pieces. A permission and embargo form, so the depositor states what may be public, what is restricted for a period, and what is metadata-only. And someone who checks the item before it is released — a five-minute review that catches the missing abstract, the wrong year and the file that is a photograph of a screen.
Step 5 — Digitise, with specifications written down
Only now, and only for material you are entitled to publish.
Scan text at 300 to 400 dpi and illustrated or rare material at 600 dpi. Keep an uncompressed master file — TIFF — untouched, and generate a delivery copy from it: PDF/A for documents, with OCR so the text is searchable. Do the OCR knowing its limits: English is reliable, Indic scripts are markedly less so, and a page of Devanagari OCR still needs a human eye before anyone relies on it.
Fix a file-naming convention on day one and never break it. Store to the 3-2-1 rule — three copies, on two kinds of media, one of them off site. An external drive in the librarian's cupboard is one copy, not a backup.
On volume: a scanned thesis runs to a few hundred megabytes as a master and a small fraction of that as a delivery PDF, so a few thousand theses is single-digit terabytes of masters and very little of what you actually serve. Storage is the cheapest part of this project. The scanner is the one place hardware money genuinely goes — an overhead book scanner costs several lakh and is worth it only if you have a real digitisation programme; an A3 flatbed handles a college's normal work for far less.
Step 6 — Remote access, before the big spending starts
A proxy such as EZproxy makes an off-campus request look as though it came from the campus once the user has logged in with college credentials. Federated access — Shibboleth, OpenAthens — is cleaner still, and lets the publisher ask the college's identity system whether this person is a member without ever seeing a password.
Whichever you choose, test it the way a student will use it: from a phone on mobile data, at night, with an account that is not yours. A subscription that only works from three machines in the reading room is not a digital library.
Step 7 — License last, and against evidence
Now buy. Use COUNTER usage reports at every renewal — publishers supply them — and be willing to cancel what nobody opens. Ask for a perpetual access clause so a cancellation does not take the years you already paid for. And check consortium routes before direct deals; the same title is frequently cheaper through one.
The recurring pattern in Indian college digitisation is not a technology failure. Koha and DSpace are free, storage is cheap and the national resources are already paid for. What is missing is a person whose job it is, and a workflow that collects material without depending on goodwill. A project with both succeeds on a small budget; a project with neither fails on a large one.
What it actually costs
The software is free. The realistic lines are: hosting or a server, at the scale of a small annual subscription; an implementation partner for the first year of Koha and DSpace, if you have no systems staff; a scanner, chosen against how much you will really digitise; storage and backup, which is minor; N-LIST membership, which is minor; and staff time, which is the largest item and the one that never appears in the budget.
The number that dominates over five years is none of the above. It is journal subscriptions — a recurring cost that rises annually and, without a perpetual access clause, leaves nothing behind. Everything in steps 1 to 6 is inexpensive infrastructure that makes step 7 spend well or badly.
Why it matters for students and researchers
For a student, the visible result of a project done in this order is small and specific: the catalogue tells you what the college has, one login opens it from the hostel, and past theses in your own department are readable instead of being three copies in a cupboard.
For a researcher, the repository is where your work becomes visible to people whose institutions do not subscribe — which in India is most people. Deposit costs an afternoon.
For a college, the argument that carries a committee is usually the accreditation one: automation, e-resources and institutional output are exactly what accreditation documentation asks about, and a working repository produces that evidence as a by-product instead of as an annual scramble. That is a real benefit. It is also, quietly, the reason to do the steps in this order rather than buying a portal that demonstrates well.
Frequently asked questions
How do I create a digital library for my college?
In sequence: appoint one owner and write the deposit policy; automate the catalogue with Koha; switch on the free national resources and N-LIST; stand up a DSpace repository with a fixed metadata minimum; attach deposit to an existing compulsory process; digitise what you own outright to written specifications; configure remote access; and only then license paid e-resources, against usage data.
What software is needed for a college e-library?
Koha for cataloguing and circulation, DSpace for the institutional repository, a proxy or federated login for remote access, and scanning software with OCR. All the core pieces are open-source and free to license; what you pay for is hosting and implementation help.
What does a digital library cost to set up in India?
Less in software than most colleges expect and more in staff time than most budget for. The free components are the software and the national resources; the real spending is a scanner sized to your digitisation plans, hosting, an implementation partner in year one, and the person running it. Recurring journal subscriptions, if you buy them, dominate the five-year cost.
Should we scan books we already own?
Only if you hold the rights or the material is out of copyright or openly licensed. Owning a physical copy does not carry the right to publish a digital one. The material a college can always digitise is its own output: theses, question papers, proceedings, institutional journals and archives.
How long does it take?
The catalogue is the long pole — retro-conversion of an existing collection is months, not weeks. A repository can be running with real content in a term if deposit is attached to thesis submission. A realistic first-year target is a live catalogue, N-LIST switched on, remote access working, and one full cohort of theses deposited.
What is the best digital library solution for an engineering college?
The same stack, with two emphases: subject e-resources matter more, so consortium routes and usage-based renewal are worth real attention; and project reports and final-year theses are a substantial, genuinely unique body of material that most engineering colleges throw away every year.