Do your builds have other hard dependencies?

The last couple of posts have discussed the need to eliminate the hard dependency of absolute paths from your build scripts and code. But there are other hard dependencies you really want to avoid, if possible, when building and running your code. Here are a few, along with ideas for eliminating them.

Registry, Environment, Config files

I recommended using the registry, environment variables, and configuration files as a way around the problem of absolute paths in your build scripts. Of course, now they’re dependent on these things, instead of the files and directories in the paths. If possible, use some of the other techniques for eliminating absolute paths from your build. If not, you can tackle these in a couple of different ways. You might make the registry settings, environment variables, and config files optional, with smart, commonly used defaults if they are not available. It’s not perfect, but it’s something. Also, it probably makes sense to script the dev machine setup, so that these things can be automatically setup before starting to build and hack on a new machine.

File shares

File shares can be like absolute paths, but even worse. Because they’re not on your machine, as a developer, you may not even have the access necessary to fix the problem on your own. And if the file share goes down or it’s path is changed for any reason, you’re going to start having builds break. You can start to solve this problem by putting the files in question into a version control system and treating it like one more component of source code. If these files are on a file share because they’re very large, it’s worth using a VCS that is centralized, at least until the DVCS’s solve the current drawbacks associated with large files.

Databases

I had not experienced the issue of needing a database to perform a build until coming to Fog Creek. The build process for Kiln generates a database through a series of version changes, then, from the db, generates typed classes used for accessing the database model. I suspect this is not too uncommon. A better approach would be to generate the database from the code model, or use a third “specification” for the database that could generate both the db and the code. But if you do stick with the database, you can at least make sure that each build creates and generates it’s own database, which is cleaned up with your standard build cleaning tools. This will let you still do multiple builds in parallel, and make sure you don’t break something in the process that is hidden by having a currently working database already in place.

Web services or sites

A build process may use web services or web sites to generate reports, store build statistics, or retrieve data used in the build. In one sense, a version control system falls into this category, and is an acceptable hard dependency. Other uses of websites or web services should be as optional as possible, and not part of the core build scripts. Make sure the core build scripts work, then collect any data and generate any reports you wish to. If you need data from an external source, seriously consider setting up a process for retrieving that data and placing it in a VCS repository on a regular schedule, so that you can just rely on your VCS.

Services or other running programs

Services or other running programs on your dev machine are also hard dependencies and should be avoided in much the same was as web services and sites if they are non-standard. In the example above, where the hard dependency of a database is overcome by having the build generate it’s own database, you’d still have the dependency of a database server instance running, probably on the same machine. For this reason, having the build generate the necessary SQL to create the database schema is a better approach than using an actual database during the build process.

Services/other programs not running (virus scanners, etc.)

Additionally, your build scripts may have an implicit “requirement” that certain services not be running, the most common of which is a virus scanner. Builds can modify tons of files, and if your virus scanner is immediately opening each modified file to check for viruses your builds may become unnecessarily slow. If possible, disable these as part of the build script, or find some way to have the scanner avoid your code and built files. Whatever you do, don’t leave a note on some internal wiki page that is impossible to find that tells new developers to disable the virus scanner before building to see a 200% speedup in builds.

Hardware available and connected (other than disk/memory/cpu)

Finally, it’s also possible, depending on the software you’re building, that your build scripts expect custom hardware available and use that hardware as part of the build process. The simplest, and most acceptable use of this is some cool indicator of build status that all of your office can see or hear. But don’t forget to make sure that hardware failure there doesn’t fail the build itself, or worse yet, stop new builds from being made. For other types of hardware dependencies, try to find a way to emulate them.

All of the hard dependencies listed above can make your life miserable, or possibly just more painful than it has to be. Finding ways to break those dependencies will keep your operations running more smoothly, your developers more happy, and your thrilled at how much more work you can get done between releases.

How do I eliminate absolute paths from my code and build scripts?

In my last post, I discussed the need for your code and build scripts to avoid absolute paths. If they do use absolute paths, those paths become hard dependencies that must be satisfied for you to be able to build your product on a given machine. Obviously, you cannot eliminate all hard dependencies (you’ll need to build within the context of an operating system, for example), but the more you eliminate the easier it will be to quickly start working on the code, rather than on your system.

The need to eliminate absolute paths applies to all paths that are used, implicitly or explicitly: paths to build tools, to source files, to target locations for built files, even to standard OS components. The more of these you can replace, the more independent your build scripts will be. Let’s look at some ways we can eliminate these paths.

Use relative paths

If you want to just replace your absolute paths with relative paths, you immediately face the question of “relative to what?” The simple answer is the current directory. But each of the other options below can and will use relative paths as well (e.g. you can append a relative path to one you get from an environment variable). The major benefit to switching to relative paths is that it will often be the most straightforward change for paths to code and the resulting binaries. To make this change, first find out if the absolute path you care about is always relative to some base path. Then get that base path using one of the methods below, and append the relative path.

Use the current directory

Often the easiest path to take advantage of is the current working directory. Any language you’re using will provide you a way to either get the current working directory path, or at least execute commands as if you were at that directory in a shell. You can also change it if necessary. When a script is always run from a given directory, you can rely on having it as the current working directory in your scripts.

Use the path to the current file

It’s really nice to be able to run scripts without switching to the directory they are in, so it’s often good in build scripts to use the path to the script or code file itself when combining with relative paths. The way to do this varies from language to language and from tool to tool, but is almost always available. One way to take advantage of this is just to cache the current working directory, change it to the directory of the current file, then do the main part of your script, then change back to the cached directory. Alternatively, you can store the path in a variable and use it as needed throughout the script.

Use the PATH variable (sparingly)

All major operating systems have the concept of a PATH variable that they’ll search when trying to execute a command. The PATH variable usually isn’t very helpful when specifying specific files, only when running commands. It’s primary use is for common build tools that don’t have any other way of providing a path for execution. Changing the PATH on your development and build machines should be done as little as possible. If there are other ways of finding a program to run (other than absolute paths), use them. However, it probably makes sense to make one change to your PATH, by adding the path to a common repository of build and development tools that are used in house. You can go pretty extreme and include all of your build tools (including standard ones like Java, the .NET framework tools, etc.) in a repository like this, or just include those developed in-house. If you have a script for setting up the development environment and tools on a new machine, use it to set add any paths needed to the system-wide PATH variable.

Use environment variables

Some tools are best accessed via environment variables other than PATH that provide a path that you can use. For example, Java uses JAVA_HOME. Stuff that ships with the operating system might be available using an environment variable like SYSTEMROOT. If necessary your build scripts can set and later use environment variables in order to communicate information about the layout of code and other files on disk.

Use system calls

From certain languages, you’ll be able to make direct OS calls to methods, like System.Environment.GetFolderPath or SHGetKnownFolderIDList, that provide paths to common locations on the system. There are typically more paths available this way than through environment variables, but it may take more work to write the code to get that information.

Use the registry

Quite frequently, installers will set registry keys that indicate the location of the tools and programs installed. You can use these registry keys to get paths for your build scripts, thus allowing developers and build managers to install things to non-standard locations, which may be useful when you need to install multiple versions of certain tools.

Use configuration files

Finally, you can also use configuration files to allow developers, or the build manager, to set or override specific paths used in the build scripts or in the code.

Just as breaking dependencies can make your code cleaner, breaking dependencies in your build scripts can make them cleaner and decouple them from a specific machine or setup. This gives you more freedom in how you manage builds and releases, seams for more easily modifying your build process, and more confidence that your scripts will work in any environment.

Where do developers keep their checked-out source code?

In asking that question, I suppose I’m implying that it matters where developers keep their checked-out source. For some teams it matters a great deal - they cannot build their code without having it at a specific location, because build scripts have hard-coded, absolute paths. For other teams, it doesn’t matter at all, and every developer puts it in a location that seems right for them.

Actually, if you know where the devs on your team keep their source code, you’ve failed. But if you don’t know where they keep their source code, you’ve failed too. You should know where the code is likely to be, but also know that it need not be there.

What you really want is for the code and build scripts to work independent of location, but have a team convention for where it goes. By making sure that the code and build scripts are independent of the path, you make it easier to set up a build machine, which may have many copies of the source code, depending on how builds and releases are managed. You also make it easy for developers to have multiple enlistments, or copies of the source code, on their machine for development purposes. And if you ever need to do some one-off stuff it shouldn’t be too painful.

On the other hand, a convention for where code is placed can grease the skids when working with other developers. If you do any pair programming it’s very important that each development machine have the same directory layout for both code and built files. Even if you don’t, there will be times when you need to help another developer, or debug an issue that only occurs on someone else’s machine. Having standard locations for the code can grease the skids to doing this kind of work.

Can you build your product with one command?

Yesterday I asked what build and release managers do. I’m going to keep asking questions, and answer them to the best of my ability as I go.

In Ship It! Jared Richardson and William Gwaltney have a chapter about scripting your build. You should be able to run a one line command that fully builds your product and that works on every developer’s machine. In the chapter, they discuss the importance of scripting your build from day one, as well as different tools for making that easier.

This is a great start, but what does it mean in real life? Real life is messy. If you’re a normal developer then you probably work on multiple products and there are interesting dependencies among those products. For example, it didn’t take long before my work on the Kiln installer led me to make fixes to the FogBugz installer (a dependency). Also, any given product is made up of many components, some of which are shared - used in other products. And a given developer may not be making changes to anything except shared components. I am sure you can come up with other interesting scenarios.

And the possibilities multiply even more when you start considering different configurations of your software. What if you need to build the 32bit version, rather than the 64bit version? Maybe you need to build a debug build to get better diagnostics, or a ship build because a bug only shows up there. Kiln and FogBugz also have licensed builds, that our customers install on their servers, as well as hosted builds, where we host the products on our servers. Each of these configurations is a different “build” that I’d want to be able to create using my build script.

What you really want is a standard way to script builds within your organization. Building any given component should take a one line command. Building any product should take a similar one line command. The easiest way to do this is to have a standard build system and script in each directory of your source code that can be built. So to build your custom ui framework, you could change to the directory where the code is and type build.bat or make or whatever. The script for a bigger product works the same way, but it builds the necessary components also.

Your build script should use command line parameters for building different configurations, with intelligent defaults. Here at Fog Creek, the default build.bat should create a licensed, debug, x64 install of FogBugz or Kiln on the developers box. But the build script should also take parameters that allow me to create a hosted, release, x32 installation of FogBugz.

All in one command.

What is build and release management?

Last month I took on some new responsibilities at Fog Creek - I am now the build and release manager for our products. Since starting here a little less than a year ago, I’ve been making regular suggestions around our processes for developing and releasing our software. That, combined with a pretty heavy load on my manager, who was spending 50-75% of his time managing builds and releases, made it obvious that we needed someone whose sole responsibility was maintaining and improving our build and release scripts and processes. Once we all realized that I would be a good fit, I started the transition to this new role about a month ago.

So what do I do? Well, I see my responsibilities falling into a few major areas.

  • Monitor our daily builds and farm out any build breaks
  • Make sure the QA team has current builds for testing purposes
  • Create and deploy major and minor releases of FogBugz, Kiln, and our supporting tools (such as the Fog Creek website)
  • Improve, streamline, refactor, etc. all of our build and deployment scripts
  • Create and maintain developer and management tools related to building and releasing our software
  • Improve our company culture around builds and releases
  • Where appropriate , be a liaison between developers and the system administrators

When I worked on Microsoft Office, they had quite a large team of developers who took care of these responsibilities. Obviously, a smaller company will probably not even have a dedicated person doing these tasks - it wouldn’t make sense. But even one person development shops should take stock of where they stand in the areas I mentioned above. With a proper focus on these tasks over the last few years, I suspect Fog Creek could have delayed dedicating someone to this work. As it stands now, there is plenty to do, and I feel much busier than I did writing new features for Kiln. As I get deeper into the work and learn more about the responsibilities involved I may change my mind. Whether or not I do, I plan to share both what I currently know and what I learn through hard experience (or just from reading good books and blogs).