|By Tarak Modi||
|October 1, 2000 12:00 AM EDT||
One of the problems of highly distributed systems is figuring out how systems discover each other. After all, the whole point of having systems distributed is to allow flexible and perhaps even dynamic configurations to maximize system performance and availability. How do these distributed components of one system or multiple systems discover each other? And once they're discovered how do we allow enough flexibility, such as rediscovery, to allow their fail-safe operation?
Space-based programming may provide us with a good answer to these questions and more. In this article I'll describe what a space is and how it can be used to mitigate some of the issues mentioned above. And I've included a technique to convert an ordinary message queue into a space.
What Is a Space?
Conventional distributed tools rely on passing messages between processes (asynchronous communication) or invoking methods on remote objects (synchronous communication). A space is an extension of the asynchronous communication model in which two processes are not passing messages to one another. In fact, the processes are totally unaware of each other.
In Figure 1, Process 1 places a message into the space. Process 2, which has been waiting for this type of message, takes the message out of the space and processes it. Based on the results, it places another message into the space. Process 3, which has been waiting for this type of message, takes the message out of the space.
Following are highlights of the preceding discussion:
- The space may contain different types of messages. In fact, I used the term message for clarity. These messages are actually just "things" (the message may be an object, an XML document or anything else that the space allows to be put in it). In Figure 1 the different shapes in the space illustrate the different types of messages.
- The three processes involved have no knowledge of one another. All they know is that they put a message in a space and get a message out of the space.
- As in the message-passing scenario, we aren't limited to two processes communicating asynchronously, but rather any number of processes communicating via a common space. This allows the creation of loosely coupled systems that can be highly distributed and extremely flexible, and can provide high availability and dynamic load balancing.
Assume that passwords can't be more than four characters in length and only alphanumeric ASCII characters are used. This gives us 14,776,336 possible passwords (624). Using the brute force technique to break the password, assume that the main program breaks the input set into 16 pieces and puts each piece along with the encrypted password in the space. The password-breaking programs watch the space for such pieces and each available program immediately grabs a piece and starts working. The programs continue until no more such pieces are available or until the password has been broken. If the password is broken, the breaking program puts the solution in the space, which is picked up by the main program.
The main program then proceeds to pick up the remaining pieces, since it has already found the solution it needs. The program never knew how many password-breaking programs were available, nor did it know where they were located. The password-breaking programs had no knowledge about one another or about the main program. If there were 16 password-breaking programs available, and each one was on a separate machine, we would've had 16 machines working on breaking the password simultaneously!
No change to any configuration of the system is required to add new password-breaking programs. This is why spaces are so good for fault tolerance, load balancing and scalability.
As you can see, spaces provide an extremely powerful concept/mechanism to decouple cooperating or dependent systems. The concept of a space isn't new, however. Tuple spaces were first described in 1982 in the context of a programming language called Linda. Linda consisted of tuples, which were collections of data grouped together, and the tuple space, which was the shared blackboard from which applications could place and retrieve tuples. The concept never gained much popularity outside of academia, however. Today spaces may be an elegant solution to many of the traditional distributed computing dilemmas. In recognition of this fact, JavaSoft has created its own implementation of the space concept, JavaSpaces, and IBM has created TSpaces, which is much more functional and complex than JavaSpaces. (We won't discuss IBM's TSpaces in this article.)
We're now in a position to describe some of the key characteristics of a space:
- Spaces provide shared access: A space provides a network-accessible "shared memory" that can be accessed by many shared remote/local processes concurrently. The space handles all issues regarding concurrent access, allowing the processes to focus on the task at hand. At the very least, spaces provide processes with the ability to place and retrieve "things." Some spaces also provide the ability to read/peek at things (i.e., to get the thing without actually removing it from the space, thus allowing other processes to access it as well).
- Spaces are persistent: A space provides reliable storage for processes to place "things." These "things" may outlive the processes that created them. It also allows the dependent/cooperating processes to work together even when they have nonoverlapping life cycles, and boosts the fault tolerance and high-availability capability of distributed systems.
- Spaces are associative. Associative lookup allows processes to "find" the "things" they're interested in. As many processes may be using/sharing the same space, many different "things" may be in the space. It's important for processes to be able to get the "things" they require without having to filter out the "noise" themselves. This is possible because spaces allow processes to define filters/templates that instruct/direct the space to "find" the right "things" for that process.
JavaSoft's Implementation: JavaSpaces
JavaSpaces technology, a new realization of the tuple spaces concept described above, is an implementation that's available free from JavaSoft. JavaSpaces is built on top of another complex technology, Jini, a Java-based technology that allows any device to become network aware. Jini provides a complex yet elegant programming model that realizes the Jini team's vision of "network anything, anytime, anywhere."
The goal of JavaSpaces is to provide what might be thought of as a file system for objects. Like other JavaSoft APIs, JavaSpaces provides a simple yet powerful set of features to developers. As I see it, however, JavaSpaces has four drawbacks:
- The implementation of JavaSpaces is complex to install.
- The fact that it builds on top of Jini makes it a little too heavy, especially if there are no plans to use Jini elsewhere in the project.
- JavaSpaces relies on Java RMI, the suitability of which for highly scalable commercial applications is a topic of debate among many software gurus.
- JavaSpaces works only with serializable Java objects.
Even though commercial implementations of spaces are available in the market, there are several reasons to create your own. If you work in a start-up company, budget constraints may be a big reason. Also, the functionality offered by a commercial implementation may be too much for the job at hand. Not only may this result in a larger learning curve, it may even slow down your application due to the sheer size of the memory footprint. Finally, it's always fun to create your own implementation.
At Online Insight we decided to create our own implementation. The primary reasons for our decision were our limited set of requirements and the extremely lightweight implementation we required to achieve our scalability and performance goals.
Our requirements can be summarized as follows:
- The space must support shared access.
- The space must be persistent.
- The space must provide the ability to specify a filtering template.
- The space must allow one "thing" to be accessed by only one process/application at a time (i.e., we don't support the "read" operation).
- The space must perform and scale well under load.
- The space must be accessible to other CORBA objects.
- The space must not impose a limitation on what you can put in it (unlike JavaSpaces, for example).
- The space must not impose size limitations on what you can put in it (the underlying hardware, however, may impose a limitation).
Java Message Service
At the time we were evaluating message queuetype software specifically, Java Message Service (JMS) implementations we realized that we could build our space facility on top of one of these queues.
JMS is an API for accessing enterprise-messaging systems from Java programs. It defines a common set of enterprise-messaging concepts and facilities, and attempts to minimize the set of concepts a Java language programmer must learn to use, including enterprise-messaging products such as IBM MQSeries. JMS also strives to maximize the portability of messaging applications. It doesn't, however, address load balancing/fault tolerance, error notification, administration of the message queue or security issues. These are all message queue vendorspecific and outside the domain of the JMS.
By using message queues that expose a JMS interface, we allow ourselves the flexibility to switch vendors of message queues if we discover that the selected one doesn't meet our scalability requirements. This separation of implementation from interface is an important design pattern (see the Bridge design pattern in Design Patterns by Gamma et al., published by Addison-Wesley). Since each JMS implementation has its own unique way of getting the initial connection factory, we defined a Java interface with one method, "getConnectionFactory", which returns the initial connection factory.
Each space is configured through a properties file. One property in this file is the fully qualified name of the class that implements this interface. There is one such class for each JMS implementation supported by the space. For example, we created one class for Sun's Java Message Queue and one for Progress Software's SonicMQ. By doing this, changing the underlying message queue used by the space is simply a matter of changing the name of the Java class in the properties file for the space. Therefore, if one vendor's message queue doesn't live up to our expectations, we can quickly switch to another.
The space implementation itself is a CORBA object that has the following interface:
void write(in ByteStream blob) raises (SpaceException);
ByteStream take() raises (SpaceException);
void write_filter(in ByteStream blob, in FilterSeq f)
ByteStream take_filter(in FilterSeq f) raises (SpaceException);
ByteStream take_filter_as_string(in string f)
The type ByteStream simply evaluates to a stream of bytes. Hence, anything that can be represented as a stream of bytes, such as a CORBA object IOR, a serialized Java object or an XML document, can be stored in the space and retrieved.
Each space instance has three properties: a name, a property that indicates if this instance of the space is persistent and a property that indicates if this instance of the space allows filters. The reason there are properties to turn the persistence and filtering off is purely for performance.
Not all spaces in our application domain are required to be persistent, in which case persistence is a performance bottleneck because it involves writing out to a database or similar storage mechanism. Similarly, if filtering isn't required, it's a performance bottleneck. As mentioned above, each space is configured through a properties file,which has the property indicating the space name, the persistence status (on/off) and the filtering status (on/off) of the space.
An example of the properties file used in configuring the space is shown below:
SpaceName=MySpaceThe "SpaceName" property is the name of the space, "AllowFilter" is a boolean property where true means the space turns filter support on and "Persistent" is a boolean property where true means the space turns persistence on. "SpaceFactory" is set to the fully qualified name of the class that allows us to get the initial connection factory from the message queue. In the foregoing example, this property is set to a class that works with SonicMQ implementation.
# The factory to use to get the initial Connection Factory
During start-up each space installs itself in the CORBA Name Service using its name property as the binding name and in the CORBA Trader Service with the name, persistence and filter properties. Thus interested applications/processes can find a space by using a well-known name from the CORBA Name Service or the space properties from the CORBA Trader Service. For example, an application that wants filtering but isn't interested in persistence can indicate these requirements to the CORBA Trader Service, which will then provide the application with a list of CORBA space references that match these requirements. The application may then choose one from that list based on some further screening.
Our implementation of the space gains all its persistence and filtering capabilities from the underlying messaging queue provider. Our space is the only client of the message queue. In our implementation the only purpose the message queue serves is as a high-quality storage/retrieval mechanism that also provides filtering capabilities. We aren't relying on the queuing facilities per se.
Each method of the CORBA interface is detailed below:
- write: This method is called by an application when it wants to put a stream of bytes into the space and doesn't want to attach filtering properties to the stream.
- write_filter: This method is used by an application when it wants to put a stream of bytes into the space and wants to attach filtering properties to the stream. The type FilterSeq evaluates to an array of filters that are attached to that bytestream. A filter is a name-value pair. Hence, a FilterSeq is an array of name value pairs.
- take: This method is called by an application when it wants to retrieve a stream of bytes from the space. No filtering is performed since none is specified.
- take_filter: This method is called by an application when it wants to retrieve a stream of bytes from the space. However, in this case a FilterSeq is provided. For a match to occur, the bytestream must have a subset of the filters provided in the method call, and the value of each filter attached to the bytestream must match the value for the corresponding filter in the method call.
- take_filter_as_string: This method is called by an application when it wants to retrieve a stream of bytes from the space. In this case a string that specifies the exact filter is provided. For a match to occur, the filter properties attached to the bytestream must satisfy the filter string provided in the method call. This method is used when the filtering conditions can't be specified as a FilterSeq.
- shutdown: This method is called to shut down the space. The shutdown is clean, which means the registration with the Name Service and the Trader Service is removed.
Distributed applications can be notoriously difficult to design, build and debug. The distributed environment introduces many complexities that aren't present when writing stand-alone applications. Some of these challenges are network latency, synchronization and concurrency, and partial failure.
Space-based programming, although not a silver bullet, is an excellent concept that can lead to an elegant solution to these problems. It takes us one step closer to achieving our goals in a distributed system, namely those of scalability, high availability, loose coupling and performance. It also helps us face the challenges mentioned above. Best of all, you don't have to buy an expensive implementation to get started with this excellent concept. It's fairly easy to create a homegrown implementation that satisfies your requirements...and it's fun, too!
- Linda Group: www.cs.yale.edu/HTML/YALE/CS/Linda/linda.html
- JavaSpaces homepage: www.javasoft.com/products/javaspaces/
- IBM, TSpaces: www.almaden.ibm.com/cs/TSpaces/
- Carriero, N.J. (1987). "Implementation of Tuple Space Machines," PhD thesis, Yale University, Department of Computer Science.
- Segall, E.J. (1993). "Tuple Space Operations: Multiple-Key Search, Online Matching and Wait-Free Synchronization," PhD thesis, Rutgers University, Department of Computer Science.
- Gul, A., et al. "ActorSpaces: An Open Distributed Programming Paradigm," University of Illinois at Urbana-Champaign, ULIUENG-92-1846.
SYS-CON Events announced today that HPM Networks will exhibit at the 17th International Cloud Expo®, which will take place on November 3–5, 2015, at the Santa Clara Convention Center in Santa Clara, CA. For 20 years, HPM Networks has been integrating technology solutions that solve complex business challenges. HPM Networks has designed solutions for both SMB and enterprise customers throughout the San Francisco Bay Area.
Aug. 1, 2015 04:45 PM EDT Reads: 481
For IoT to grow as quickly as analyst firms’ project, a lot is going to fall on developers to quickly bring applications to market. But the lack of a standard development platform threatens to slow growth and make application development more time consuming and costly, much like we’ve seen in the mobile space. In his session at @ThingsExpo, Mike Weiner, Product Manager of the Omega DevCloud with KORE Telematics Inc., discussed the evolving requirements for developers as IoT matures and conducted a live demonstration of how quickly application development can happen when the need to comply wit...
Aug. 1, 2015 03:15 PM EDT Reads: 326
The Internet of Everything (IoE) brings together people, process, data and things to make networked connections more relevant and valuable than ever before – transforming information into knowledge and knowledge into wisdom. IoE creates new capabilities, richer experiences, and unprecedented opportunities to improve business and government operations, decision making and mission support capabilities.
Aug. 1, 2015 10:00 AM EDT Reads: 294
Explosive growth in connected devices. Enormous amounts of data for collection and analysis. Critical use of data for split-second decision making and actionable information. All three are factors in making the Internet of Things a reality. Yet, any one factor would have an IT organization pondering its infrastructure strategy. How should your organization enhance its IT framework to enable an Internet of Things implementation? In his session at @ThingsExpo, James Kirkland, Red Hat's Chief Architect for the Internet of Things and Intelligent Systems, described how to revolutionize your archit...
Jul. 30, 2015 07:30 PM EDT Reads: 1,413
MuleSoft has announced the findings of its 2015 Connectivity Benchmark Report on the adoption and business impact of APIs. The findings suggest traditional businesses are quickly evolving into "composable enterprises" built out of hundreds of connected software services, applications and devices. Most are embracing the Internet of Things (IoT) and microservices technologies like Docker. A majority are integrating wearables, like smart watches, and more than half plan to generate revenue with APIs within the next year.
Jul. 30, 2015 02:30 PM EDT Reads: 124
Growth hacking is common for startups to make unheard-of progress in building their business. Career Hacks can help Geek Girls and those who support them (yes, that's you too, Dad!) to excel in this typically male-dominated world. Get ready to learn the facts: Is there a bias against women in the tech / developer communities? Why are women 50% of the workforce, but hold only 24% of the STEM or IT positions? Some beginnings of what to do about it! In her Opening Keynote at 16th Cloud Expo, Sandy Carter, IBM General Manager Cloud Ecosystem and Developers, and a Social Business Evangelist, d...
Jul. 30, 2015 12:00 PM EDT Reads: 2,068
In his keynote at 16th Cloud Expo, Rodney Rogers, CEO of Virtustream, discussed the evolution of the company from inception to its recent acquisition by EMC – including personal insights, lessons learned (and some WTF moments) along the way. Learn how Virtustream’s unique approach of combining the economics and elasticity of the consumer cloud model with proper performance, application automation and security into a platform became a breakout success with enterprise customers and a natural fit for the EMC Federation.
Jul. 30, 2015 09:00 AM EDT Reads: 2,170
The Internet of Things is not only adding billions of sensors and billions of terabytes to the Internet. It is also forcing a fundamental change in the way we envision Information Technology. For the first time, more data is being created by devices at the edge of the Internet rather than from centralized systems. What does this mean for today's IT professional? In this Power Panel at @ThingsExpo, moderated by Conference Chair Roger Strukhoff, panelists addressed this very serious issue of profound change in the industry.
Jul. 29, 2015 03:00 PM EDT Reads: 1,287
Discussions about cloud computing are evolving into discussions about enterprise IT in general. As enterprises increasingly migrate toward their own unique clouds, new issues such as the use of containers and microservices emerge to keep things interesting. In this Power Panel at 16th Cloud Expo, moderated by Conference Chair Roger Strukhoff, panelists addressed the state of cloud computing today, and what enterprise IT professionals need to know about how the latest topics and trends affect their organization.
Jul. 29, 2015 02:00 PM EDT Reads: 1,194
It is one thing to build single industrial IoT applications, but what will it take to build the Smart Cities and truly society-changing applications of the future? The technology won’t be the problem, it will be the number of parties that need to work together and be aligned in their motivation to succeed. In his session at @ThingsExpo, Jason Mondanaro, Director, Product Management at Metanga, discussed how you can plan to cooperate, partner, and form lasting all-star teams to change the world and it starts with business models and monetization strategies.
Jul. 28, 2015 04:30 PM EDT Reads: 1,772
Converging digital disruptions is creating a major sea change - Cisco calls this the Internet of Everything (IoE). IoE is the network connection of People, Process, Data and Things, fueled by Cloud, Mobile, Social, Analytics and Security, and it represents a $19Trillion value-at-stake over the next 10 years. In her keynote at @ThingsExpo, Manjula Talreja, VP of Cisco Consulting Services, discussed IoE and the enormous opportunities it provides to public and private firms alike. She will share what businesses must do to thrive in the IoE economy, citing examples from several industry sectors.
Jul. 28, 2015 11:00 AM EDT Reads: 2,047
There will be 150 billion connected devices by 2020. New digital businesses have already disrupted value chains across every industry. APIs are at the center of the digital business. You need to understand what assets you have that can be exposed digitally, what their digital value chain is, and how to create an effective business model around that value chain to compete in this economy. No enterprise can be complacent and not engage in the digital economy. Learn how to be the disruptor and not the disruptee.
Jul. 27, 2015 10:00 AM EDT Reads: 2,038
Akana has released Envision, an enhanced API analytics platform that helps enterprises mine critical insights across their digital eco-systems, understand their customers and partners and offer value-added personalized services. “In today’s digital economy, data-driven insights are proving to be a key differentiator for businesses. Understanding the data that is being tunneled through their APIs and how it can be used to optimize their business and operations is of paramount importance,” said Alistair Farquharson, CTO of Akana.
Jul. 27, 2015 09:00 AM EDT Reads: 330
Business as usual for IT is evolving into a "Make or Buy" decision on a service-by-service conversation with input from the LOBs. How does your organization move forward with cloud? In his general session at 16th Cloud Expo, Paul Maravei, Regional Sales Manager, Hybrid Cloud and Managed Services at Cisco, discusses how Cisco and its partners offer a market-leading portfolio and ecosystem of cloud infrastructure and application services that allow you to uniquely and securely combine cloud business applications and services across multiple cloud delivery models.
Jul. 27, 2015 08:00 AM EDT Reads: 1,907
The enterprise market will drive IoT device adoption over the next five years. In his session at @ThingsExpo, John Greenough, an analyst at BI Intelligence, division of Business Insider, analyzed how companies will adopt IoT products and the associated cost of adopting those products. John Greenough is the lead analyst covering the Internet of Things for BI Intelligence- Business Insider’s paid research service. Numerous IoT companies have cited his analysis of the IoT. Prior to joining BI Intelligence, he worked analyzing bank technology for Corporate Insight and The Clearing House Payment...
Jul. 26, 2015 09:00 PM EDT Reads: 1,581
"Optimal Design is a technology integration and product development firm that specializes in connecting devices to the cloud," stated Joe Wascow, Co-Founder & CMO of Optimal Design, in this SYS-CON.tv interview at @ThingsExpo, held June 9-11, 2015, at the Javits Center in New York City.
Jul. 25, 2015 02:00 PM EDT Reads: 399
SYS-CON Events announced today that CommVault has been named “Bronze Sponsor” of SYS-CON's 17th International Cloud Expo®, which will take place on November 3–5, 2015, at the Santa Clara Convention Center in Santa Clara, CA. A singular vision – a belief in a better way to address current and future data management needs – guides CommVault in the development of Singular Information Management® solutions for high-performance data protection, universal availability and simplified management of data on complex storage networks. CommVault's exclusive single-platform architecture gives companies unp...
Jul. 25, 2015 01:00 PM EDT Reads: 1,965
Electric Cloud and Arynga have announced a product integration partnership that will bring Continuous Delivery solutions to the automotive Internet-of-Things (IoT) market. The joint solution will help automotive manufacturers, OEMs and system integrators adopt DevOps automation and Continuous Delivery practices that reduce software build and release cycle times within the complex and specific parameters of embedded and IoT software systems.
Jul. 25, 2015 12:15 PM EDT Reads: 481
"ciqada is a combined platform of hardware modules and server products that lets people take their existing devices or new devices and lets them be accessible over the Internet for their users," noted Geoff Engelstein of ciqada, a division of Mars International, in this SYS-CON.tv interview at @ThingsExpo, held June 9-11, 2015, at the Javits Center in New York City.
Jul. 25, 2015 12:00 PM EDT Reads: 1,544
Internet of Things is moving from being a hype to a reality. Experts estimate that internet connected cars will grow to 152 million, while over 100 million internet connected wireless light bulbs and lamps will be operational by 2020. These and many other intriguing statistics highlight the importance of Internet powered devices and how market penetration is going to multiply many times over in the next few years.
Jul. 25, 2015 09:00 AM EDT Reads: 1,494