Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

GraphX is Apache Spark’s API for graphs and graph-parallel computation. To use it, represent your network as a Graph[VD, ED], where vertices and edges carry properties, then apply graph operators or algorithms such as PageRank and connected components. This hands-on introduction uses Scala and the Spark 3.5.7 programming guide; check the documentation for your installed Spark version because APIs and setup details can vary.

What GraphX represents

GraphX extends Spark’s RDD programming model with a distributed, immutable property graph. A Graph[VD, ED] is a directed multigraph: VD is the type of each vertex’s property, and ED is the type of each edge’s property. Each vertex has a unique 64-bit ID, called a VertexId. Because the graph is directed, an edge from one ID to another does not automatically mean the reverse relationship exists; because it is a multigraph, multiple edges between the same pair of vertices are possible.

For example, a social network might store a user’s name or account attributes as the vertex property and use directed edges to represent “follows.” The edge direction matters: following from user A to user B is not the same as B following A. See the Spark 3.5.7 GraphX Programming Guide for the version-specific API details.

Load an edge list or build a graph

Load edges from a file

GraphX provides GraphLoader.edgeListFile for edge-list files containing source and destination vertex IDs. Lines beginning with # are treated as comments and skipped. This is a minimal Scala example; set path to a file accessible to Spark:

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import org.apache.spark.graphx._
import org.apache.spark.rdd.RDD

val path = "data/edges.txt"
val graph = GraphLoader.edgeListFile(sc, path)

The loader constructs a graph from the listed relationships. If you need meaningful vertex properties, join in a separate vertex dataset or construct a graph with explicit vertex and edge RDDs. The guide’s loader and graph-construction examples are documented for Spark 3.5.7; confirm the matching guide for your deployed release.

Construct vertices and edges explicitly

When creating a graph from data already in RDDs, an edge has a source ID, destination ID, and property. The following small example stores names on vertices and relationship labels on edges:

val users: RDD[(VertexId, String)] = sc.parallelize(Seq(
  (1L, "Ada"),
  (2L, "Linus"),
  (3L, "Grace")
))

val relationships: RDD[Edge[String]] = sc.parallelize(Seq(
  Edge(1L, 2L, "follows"),
  Edge(2L, 3L, "follows"),
  Edge(1L, 3L, "follows")
))

val graph = Graph(users, relationships, "unknown")

The default vertex value, here "unknown", is used for edge endpoints that do not appear in the supplied vertex RDD. Choose a default that makes sense for your data rather than silently treating missing vertex records as complete records.

Transform a graph and aggregate neighbor data

Filter with subgraph

GraphX graph transformations return new graph values rather than modifying the original. For instance, filter out vertices that do not meet a condition with subgraph:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
val knownUsers = graph.subgraph(vpred = (_, name) => name != "unknown")

This produces a graph restricted by the vertex predicate. Use predicates on edges as well when the relationships themselves must satisfy a condition.

Combine data with joinVertices

joinVertices updates vertex properties by joining in an RDD keyed by vertex ID. For example, suppose scores contains a numeric score for some users:

val scores: RDD[(VertexId, Double)] = sc.parallelize(Seq(
  (1L, 0.8),
  (3L, 0.4)
))

val scoredGraph = graph.joinVertices(scores) { (_, name, score) =>
  (name, score)
}

The resulting graph has a tuple as its vertex property for vertices with matching score records. If you need an explicit value for vertices without a match, handle that case in your data preparation or use an appropriate outer-join workflow.

Use aggregateMessages for neighborhood summaries

aggregateMessages lets each edge context send a message to one or both endpoint vertices, then combines messages by destination vertex ID. The example below sends each source vertex’s out-degree to its destination and sums those values:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
val incomingSourceDegree = graph.aggregateMessages[Int](
  triplet => triplet.sendToDst(triplet.srcAttr match {
    case _ => 1
  }),
  _ + _
)

This illustrative aggregation sends a constant value per incoming edge, so it counts incoming edges; if the intended message depends on an attribute such as a source score, read that property from triplet.srcAttr and use the corresponding type. For performance, the guide recommends constant-sized messages and aggregations—such as numbers combined with addition—rather than repeatedly concatenating lists. That is a design guideline, not a guarantee of a particular runtime.

Choose a built-in graph algorithm

Algorithm Question it answers Important choice or condition
PageRank Which vertices are relatively important under a link or endorsement interpretation? Use a fixed iteration count for a bounded run, or a convergence tolerance for the convergence-based form.
Connected components Which vertices belong to the same connected component? GraphX labels each component with its lowest-numbered vertex ID.
Triangle counting How many triangles pass through each vertex, as a clustering signal? Edges must be oriented canonically with srcId < dstId, and the graph should be partitioned with Graph.partitionBy.

The official Spark guide documents these algorithms and their operational choices. GraphX’s project page also lists label propagation, strongly connected components, and SVD++ among its algorithm library offerings: Apache Spark GraphX.

Run PageRank

For a fixed number of iterations, GraphX offers PageRank.run; for a convergence-based run, use PageRank.runUntilConvergence. For example, a bounded run can be written as:

val rankedGraph = graph.pageRank(0.0001, 10)

Here the first argument is the reset probability (often called the PageRank tolerance parameter in the GraphX API) and the second is the iteration count for this API form; if you want to specify the usual damping factor directly, consult the version’s method documentation rather than assuming the parameter has that meaning. Choose the algorithm form and parameters based on whether a fixed compute budget or convergence behavior is the goal.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Prepare triangle counting correctly

Triangle counting is not a drop-in call for arbitrary directed edges. Before invoking it, make sure each undirected relationship has a consistent canonical orientation so the source ID is less than the destination ID, and partition the graph with Graph.partitionBy. The GraphX guide documents this precondition because the algorithm relies on that representation.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Use Pregel for iterative custom computation

GraphX’s Pregel variant expresses iterative graph computation in supersteps. Each active vertex receives inbound messages and updates its state; a user-defined send function emits messages along graph edges. The process ends when there are no messages left to send or the configured iteration limit is reached. The current Spark 4.2.0 GraphX ScalaDoc describes the Pregel API; compare it with the documentation for the version you run.

At a high level, an invocation supplies an initial message, a vertex-program function, a message-sending function, and a message merge function:

val result = graph.pregel(initialMessage, maxIterations = 10)(
  vprog = (id, state, message) => updateState(state, message),
  sendMsg = triplet => messagesForNeighbors(triplet),
  mergeMsg = (left, right) => combine(left, right)
)

The function names above describe the roles you implement; replace them with functions appropriate to your computation. Keep messages and merged state compact where possible, and set a sensible iteration limit so a non-converging computation has a bound.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Keep reused graphs and iterative jobs manageable

Cache graphs reused across actions

A GraphX value is not automatically persisted simply because it is a graph. If multiple actions reuse it, call cache() so Spark can avoid recomputing its underlying data when possible:

val reusable = graph.cache()

Caching consumes cluster memory (and may spill depending on storage behavior), so persist graphs that are actually reused and release them when no longer needed.

Use Pregel and checkpointing thoughtfully

For iterative algorithms, the Spark 3.5.7 guide recommends Pregel because it handles unpersisting intermediate results. Long lineage chains can also lead to stack overflow. For suitable long-running computations, configure a checkpoint directory and set spark.graphx.pregel.checkpointInterval to a positive interval so lineage can be truncated. This is tuning guidance for deep iterations, not required setup for a short example.

Run GraphX locally or on a cluster

The Spark project describes GraphX as a Spark module that can run locally on a multicore machine or in distributed cluster mode. Use the Spark installation and deployment method that fits your environment, and consult the documentation matching that version. The project page’s release listing is time-sensitive; verify current releases on the GraphX project page or Spark downloads before choosing a version.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.