File Directory Structure

Introduction

Every computer user who uses files and directories has the same problem: Structuring such files and directories in a good way.

The structure should help to identify where to store new files, where to find existing files and distinguish important and less important files.

This document identifies some common approaches and describes what works acceptably well at the Ma_Sys.ma. No ultimate solution to this problem has been found yet and it seems that in general, the optimal choice of organization heavily depends on the person, too.

General Rules

The following rules have been found to work across a wide range of file and directory organization contexts:

  1. Do not nest too deeply. Deep nesting makes important files hard to find. Navigating deep structures is usually slower than flat structures. How far to take this rule depends on the program under use. Some programs allow efficiently dealing with large numbers of subdirectories next to each other whereas other programs perform better in the presence of only a few subdirectories at once.

  2. Use a common subtree for important files. Many operating systems mix program generated files with user files in the home directory. Mixing cache, trash and important files is bad for backups. Backing up large trees takes longer and is thus more rarely done. Thus, important files should be below a dedicated subtree which is not mixed with less important files. Also, it helps to not spread files all over the system (or network drive). Prefer to use one subtree for all stuff that is worth backing up.

  3. Use similarly deep nesting for similarly important files. Sometimes there is a chaos where e.g. a document is put near the upper part of a hierarchy and other, related but equally important documents are placed further down. Structures are easier to navigate if files and directories are not mixed with each other. It can still help to have a few files like README.txt or important results in the top-levels of directories because such files might help short-cutting the navigation. Example: If you work on a product with types A and B which has common, type-A-only and type-B-only documents, the recommended general structure is as follows:

    /product
     |
     +-- common/
     |    |
     |    +-- doc1.txt
     |    |
     |    +-- doc2.txt
     |
     +-- type-a/
     |    |
     |    +-- doc3.txt
     |    |
     |    +-- doc4.txt
     |
     +-- type-b/
          |
          +-- ...

    Do *not* make such a structure:

    /product
     |
     +-- doc1.txt
     |
     +-- doc2.txt
     |
     +-- type-a/
     |    |
     |    +-- doc3.txt
     |    |
     |    +-- doc4.txt
     |
     +-- type-b/
          |
          +-- ...
  1. Use only safe characters in file and directory names. Some scripts and Makefile have trouble to process spaces in file names. Some programs have trouble to process Umlaute (äöüß) in file names. Avoid all problematic characters. Use characters from this set: [a-zA-Z0-9_.+-]. Prefer lowercase file names if possible. Decide deliberately when to use - or _ and prefer one of these characters to use by default. The Ma_Sys.ma preference is to prefer _ but valid arguments can be found to prefer - instead.

  2. Tags are chaos. There are some people who claim that one ought to ditch the hierarchical file structure in favor of using tags. Any file could have multiple tags and finding stuff would work by tag. The problem with tags is that tags generally lack the spatial notion of file locations. Without setting up a graph of tags, it becomes impossible to find related stuff. Also, maintaining tags is about as hard as maintaining a directory hierarchy. Hence from Ma_Sys.ma point of view, tags are not the solution. In the rare cases, that a given file needs to be accessible from multiple locations, it seems better to create a symlink (e.g. from tax filings to an invoice) vs. using tags.

Organization Approaches

Any organization may show repeating patterns. Like with Design Patterns in Object Oriented Programming, It helps to be aware of the existence of patterns because it may help to consciously decide for a pattern knowing some of its general advantages or disadvantages.

Organization by Time

Organization by Topic

Organization by Type

Organization by Lifecycle

Organization by Ticket

Hamster System

Summary

The hamster system puts the Organization by Lifecycle at the top-level of the structure. Here are two links which describe systems of that syle:

Specific implementations may vary, but in general there is the idea that on top level you have two directories:

/
 |
 +-- working/
 |
 +-- archive/

The archive may be organized by topic (up to two levels) or by time (one level).

The Plaintext Productivity approach suggests to create the permanent location for each project right in archive from the beginning and only link to it from the working directory. This reduces the disadvantage that organizing by lifecycle requires moving the directory on lifecycle change as only the link is deleted.

Comments

Both links given are worth reading because those systems are defined quite succinctly and may offer a solution for many use cases without much overhead.

For Ma_Sys.ma use cases, the distinction in working and archive on the top-level feels incomplete, because there is a lot of rarely changed files which are not archived but rather only “rarely needed” and it feels unnatural that subsequent changes would go some long-past year directory. On the other hand, keeping all rarely needed files in the “working” directory doesn’t seem to be the point of the system, either…

PARA

I haven’t checked in detail, but there is a system called PARA (Project, Area, Resource, Archive) which seems to resolve my main complaint about the working/archive distinction of the Hamster system by adding “area” for long-term responsibility topics and “resources” for documentation and stuff that may be useful later.

Links:

When starting from zero, it may be a worthwhile approach to check, because it looks like it is more elaborate compared to the Hamster system but still less complex than the Johnny Decimal system.

Johnny Decimal

Summary

The Johnny Decimal (JD) system is described at https://johnnydecimal.com/. Since its inception, the documentation about it has come to be quite extensive.

This section gives a short summary. JD is a combination of the Organization by Topic and Organization by Time with a focus on limiting the conceptual nesting depth.

It achieves this by assigning each directory an ID, either of form AC.ID or SYS.AC.ID.

The definition of these acronyms is as follows:

SYS System range A00 .. Z99
A Area range 0..9
C Category range 0..9
ID Running counter range 00..99

Ignoring the SYS part for now, the idea is that on the top level, there are at most 10 areas and inside each area at most 10 categories which form the AC part. E.g. Area 1 and Category 4 is AC=41. Inside the category, new directories are created by assigning a running counter as ID.

JD suggests to materialize the area as its own directory and creating the categories as subdirectories of it as follows:

/
 |
 +-- 10-19 My Area/
      |
      +-- 11 My First Category/
      |    |
      |    +-- 11.01 My First Directory/
      |    |     |
      |    |     +-- document.txt
      |    |
      |    +-- 11.02 My Second Directory/
      |
      +-- 12 My Second Category/
           |
           +-- 12.01 My Other First Directory/

Above tree deliberately contains spaces as the original JD system is defined with them.

The SYS numbering extends this approach by providing multiple such namespaces. E.g. a suggested distinction could be to create one number space for files at work and one for files home.

The numbering starting at 11 (rather than 10) and 01 rather than 00 for AC/ID is deliberate and reserves the special values for special purposes which may of course be defined by the user of the system.

Comments

For Ma_Sys.ma use cases, the use of spaces in file names is advised against, but the system can be easily adjusted (e.g. 10t19_administrative or even 10_administrative in favor of 10-19 Administrative work just as well).

Also, materializing the areas as distinctive directories seems to be wasteful and leans towards nesting too deeply.

At the Ma_Sys.ma, the AC are directly placed at the top-level leading to the existence of currently ~40 top-level directories. They can be selected efficiently with VIFM and (including some headings for the areas) fit well on three screens when using a 25 character high terminal.

The SYS numbering in front doesn’t make all that much sense. For Ma_Sys.ma use cases, it seems like only 2 or 3 “systems” will ever be needed (personal, work and maybe a third tree in the future). It doesn’t make sense to reserve more than one character for the system-level.

An earlier version of JD proposed PRJ instead of SYS to distinguish projects. It might work for freelance workers who could create entire trees for each large project and reserve some PRJ numbers in the beginning for their personal structure.

On the other hand, at the Ma_Sys.ma there are tons of small software projects e.g. the GIT repositories for MDVL which is locally 145 directories. For Ma_Sys.ma use cases it was thus found most helpful to allow selected areas to extend the ID counter to three digits, i.e. the Ma_Sys.ma numbering for some subdirectories is AC.III with III being the three digit ID in range 000..999.

A downside of the plain JD system is also that it doesn’t at all account for lifecycle concerns with the general idea being that lifecycle is mostly implicit in the organization by time from starting with ID=10 and incrementing onwards.

For the Ma_Sys.ma use case, selected categories have been linked together as “lifecycle structures” e.g. as follows:

/
 |
 +-- 34_items_research/
 |
 +-- 35_items_active/
 |    |
 |    +-- 35.457_sandisk_32g_usb_pendrive/
 |
 +-- 36_items_tx_complete/
 |
 +-- 37_items_end_of_life/

This is a slightly simplified visualization of the area 30 administrative buy which contains subdirectories for bought items, e.g. the USB pendrive is currently in working condition hence under active and it may move to end_of_life once it becomes defective or to tx_complete if handed over to someone else (tx_complete is shorthand for transaction complete and marks any item which was DoA, refunded, gifted or is otherwise not exactly EoL).

Personal File Structure

This is a slightly simplified and redacted version of the Ma_Sys.ma personal file structure. The directories with triple underscores are used to label the area and VIFM is configured to highlight them using a particular background color for extra structural visibility.

10___ADMINISTRATIVE_IMPORTANT___
12_keys_now/            \_ lifecycle structure. key files, ID card, ...
13_keys_old/            /
14_email/               subtree for e-mail program usage
15_important/           e.g. documentation of backup locations

20___ADMINISTRATIVE_NO_BUY___
20_apartment/           notes about living
24_tax/                 tax docs (one subdirectory per year)
26_finance_now/         \_ lifecycle structure. banks, credit card...
27_finance_old/         /
28_contracts_now/       \_ lifecycle structure. insurance, phone, ...
29_contracts_old/       /

30___ADMINISTRATIVE_BUY___
31_car/                 Car-related stuff gets its own separate subtree.
34_items_research/      \__ lifecycle structure about bought or
35_items_active/        |   planned to buy items including IT stuff
36_items_tx_complete/   |
37_items_end_of_life/   /

40___MAINQ___
41_school/
42_internships/
43_university/
45_realwork/            Current work contract with the intention to
                        create a new 46_work directory if I change
                        employer. I don't store files *from* work at
                        home but there is a lot of metadata like pay
                        slips, CV, contract etc.

50___SIDEQ___
51_family/              Subdirectories for family members
52_...                  Locations for nonprofit related activities

60___PERSONAL___
61_pictures/            Personal photos
65_gam_cnt/             Game save files
66_doc_individual/      one-off documents
67_doc_recurring/       recurring documents like calendar, birthday, ...
69_comms_excerpts/      Chat logs and such

70___DEVELOPMENT___     All directories in these area are considered
                        part of the same lifecycle structure. It allows
                        starting development on something and later
                        moving it to the MDVL subtree for publishing on
                        Github or package building.
72 dev_active/          Development which is not (yet) part of MDVL
73_dev_mdz/             MDVL subtree
74_dev_completed/       successfully completed (not on Github)
75_dev_attic/           not completed, but also not pursued further

80___KNOWLEDGE_BASE___
81_man_d5i/             D5Man legacy files
82_man_rr/              D5Man2 files
83_archivebox/          Archivebox and IAL structure
84_doc_cpy/             Scanned documents and stuff which is not in IAL

90___SYSTEM___
91_webbrowsing/         Web browser profile directory.

00_waystone.sh          Mad script which prints for each lifecycle
                        structure the max. ID which is currently
                        assigned in any of its associated directories.
                        Hardcodes the structure shown above :)

This is already the second attempt to setup JD at the Ma_Sys.ma. It works OK-ish but not perfect, e.g. some substructures are still quite a mess.

In general, the trick seems to be to deviate from the JD system in the places where is is really valuable, but stick to it as a reasonable default.

Work File Structure

The work file structure is smaller because there are less work-related files to structure compared to personal files. This is obviously because work is a team effort and most important files are ultimately going to end up in a work system whose structure (e.g. repository, network drive etc.) is managed outside of my personal namespace.

10t19_administrative    Time tracking, Form delivery, Work note files.
                        Mandatory recurring training materials. Key
                        files. Useful code snippets, scripts.
30t39_tickets           One category for each ticket system.
                        Inside that, one directory per ticket.
80t89_courses           External education/conferences/seminars etc.
                        One subdirectory per origin i.e. if the same
                        course is offered next year, it does not get an
                        own new top-level entry in here.
90t99_checkouts         Checked out repositories and large software
                        trees. The area is further distinguished into
                        work-related repositories and external (e.g.
                        free software) repositories.

There is also some 20t29 and 40t49 trees related to some one-off stuff e.g. small tests or business travels which is neither tracked in a ticket nor a repository but it is hard to give a coherent description here. It shows that despite my best efforts to setup a reasonable structure, JD is not trivial to get right from the beginning.

Conclusion

The primarily static JD assignment which models lifecycle states only in selected places (and as an out-of-system extension) and the rigid limiting of subdirectory nesting depth work very well for the cases where the structure is readily established and tested.

It needs some courage to occasionally kill some old numbering by restructuring it – what might sound like heresy to a JD maximalist is occasionally needed in any long-term deployment.

Future Directions

JD works well enough that for now it does not seem worthwhile to investigate the other systems more deeply. Instead, it is expected that at least for the personal structure, changes will only be made to selected categories and thus, in an incremental fashion.


Ma_Sys.ma Website 5 (1.0.3) – no Flash, no JavaScript, no Webfont, no Copy Protection, no Mobile First. No bullshit. No GUI needed. Works with any browser.

Created: 2026/08/29 16:17:57 | Revised: 2026/08/29 21:37:07 | Tags: kb, filesystem, file, directory, structuring, hierarchy, fhs, johnny, jd, decimal | Version: 1.0.0 | SRC (Pandoc MD) | GPL

(c) 2026 Ma_Sys.ma info@masysma.net.

This program is free software: you can redistribute it and/or modify it under the terms of the GNU General Public License as published by the Free Software Foundation, either version 3 of the License, or (at your option) any later version.
This program is distributed in the hope that it will be useful, but WITHOUT ANY WARRANTY; without even the implied warranty of MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the GNU General Public License for more details.
You should have received a copy of the GNU General Public License along with this program. If not, see <http://www.gnu.org/licenses/>.